Logo for PowerPlan, Inc.

Principal Site Reliability Engineer

Role overview

Qualifications

  • Deep hands-on experience operating production systems in AWS and Azure environments
  • Strong automation skills using Python and PowerShell in operational contexts
  • Proven ability to identify repetitive operational work and eliminate it through automation
  • Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring

Responsibilities

  • Resolve escalated infrastructure cases across major AWS and Azure services
  • Eliminate or significantly reduce manual intervention for the top 5–7 highest-frequency operational issues
  • Establish a consistent, high-quality incident response and post-incident review process
  • Deliver a mature observability layer across AWS and Azure with service-level dashboards

About the company

PowerPlan, Inc. logo

PowerPlan, Inc.

Computer Software / SaaS

PowerPlan’s award-winning integrated platform gives key stakeholders in accounting, tax, finance, operations, IT and regulatory the clarity they need to make decisions that improve corporate performance. PowerPlan layers complex regulatory requirements with granular financial and operational data from every corner of your organization into a single source of defensible and auditable information. That’s why PowerPlan developed their integrated platform to give key stakeholders in accounting, tax, finance, operations, IT and regulatory the insight they need to make better financial decisions. More than a financial solution, PowerPlan is a powerful insight machine that enables energy companies to: •Combine granular financial and operational asset details from every corner of your organization into a unique view for each stakeholder that enables you to create optimal department strategies. •Mitigate compliance risk by applying complex tax and industry specific regulatory requirements to your consolidated, auditable set of financial books. •Develop defensible, strategic financial asset scenarios that enable optimal planning for today, tomorrow and the next 20+ years and more. For more information how PowerPlan’s solutions can help your organization gain a clearer financial picture, email clarity@powerplan.com or visit www.powerplan.com.

Company details

Company typeSME
IndustryComputer Software / SaaS
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Overview:

This is a principal-level individual contributor role at the heart of our cloud platform’s reliability, scalability, and operational maturity. You will work hands-on across AWS and Azure environments, solving complex production problems while systematically eliminating the manual toil that creates them. The role offers significant autonomy, deep technical impact, and the opportunity to shape how reliability engineering is practiced across the organization.

 

COMPANY

PowerPlan operates a growing SaaS platform supporting enterprise customers with mission-critical workloads. We run complex, multi-cloud environments and value engineers who take ownership, think in systems, and build solutions that scale. Our culture emphasizes operational excellence, blameless learning, and collaboration across Engineering, Support, Professional Services, and Product teams.

Responsibilities:

KEY PERFORMANCE OBJECTIVES (First 12 Months)

 

OBJECTIVE 1: Platform Familiarity Through Escalations & Early Automation (First 90 Days)

 

Outcome:
Within 90 days, resolve escalated infrastructure cases across major AWS and Azure services and deliver 2–3 targeted automations that measurably reduce manual resolution time for recurring issues.

Impact:
Accelerates ramp-up, demonstrates immediate value, and establishes the expectation that operational issues are systematically automated rather than repeatedly handled manually.

How:
Work directly on escalated cases from Support and Professional Services, document manual resolution steps, identify repeatable patterns, and implement focused Python or PowerShell automations tied to high-frequency workflows.

 

OBJECTIVE 2: Eliminate Top Sources of Operational Toil (3–6 Months)

 

Outcome:
Within 3–6 months, eliminate or significantly reduce manual intervention for the top 5–7 highest-frequency operational issues through automation, self-service tooling, or infrastructure improvements.

Impact:
Reduces support load, improves service stability, and frees Cloud Engineering capacity for higher-value reliability and platform initiatives.

How:
Analyze case and incident data, prioritize automation candidates by frequency and impact, build production-grade automations and runbooks, and partner with Support and PS teams to validate adoption and effectiveness.

 

OBJECTIVE 3: Mature Incident Response & Post‑Incident Learning (6–9 Months)

 

Outcome:
By month 9, establish a consistent, high-quality incident response and post-incident review process resulting in faster containment, clearer ownership, and tracked corrective actions for all critical production incidents.

Impact:
Reduces repeat incidents, improves on-call effectiveness, and increases organizational confidence during high-severity events.

How:
Lead critical incidents, standardize incident runbooks, facilitate blameless postmortems, track follow-up actions to completion, and coach teams on effective incident communication and decision-making.

 

OBJECTIVE 4: Deliver a Mature, SLO‑Aligned Observability Platform (9–12 Months)

 

Outcome:
By month 12, deliver a mature observability layer across AWS and Azure with service-level dashboards, tuned alerts, and clear SLI/SLO reporting actively used by on-call and engineering teams.

Impact:
Improves detection, diagnosis, and prevention of production issues while reducing alert fatigue and enabling data-driven reliability decisions.

How:
Design Grafana dashboards aligned to service health and user journeys, integrate metrics, logs, and traces from core platforms, tune alert thresholds, and embed observability into CI/CD and incident workflows.

Qualifications:

WHAT YOU BRING

  • Deep hands-on experience operating production systems in AWS and Azure environments
  • Strong automation skills using Python and PowerShell in operational contexts
  • Proven ability to identify repetitive operational work and eliminate it through automation
  • Experience leading incident response and blameless post-incident reviews
  • Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring
  • Ability to influence engineering practices without formal authority
  • Clear written and verbal communication skills across technical and non-technical audiences

PowerPlan is an EOE

Applicant and Candidate Privacy Notice

 

Please note that this is a hybrid role that involves a combination of onsite work from our corporate office as well as work from home. While we strive to accommodate flexible working arrangements when sensible, there will be times when onsite work is required. This could include scheduled office days, team meetings, client meetings, or special events.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at PowerPlan, Inc.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.