Logo for Electric Power Engineers

Senior Site Reliability Engineer

Role overview

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related discipline
  • Strong hands-on experience designing and operating production workloads in AWS
  • Advanced proficiency with Terraform
  • Strong Linux systems knowledge and ability to diagnose issues

Responsibilities

  • Define, measure, and improve service reliability using SLIs, SLOs, error budgets, and capacity planning
  • Build and maintain monitoring, logging, tracing, dashboards, and alerting
  • Participate in and improve production incident response
  • Optimize infrastructure for reliability and performance

About the company

Electric Power Engineers logo

Electric Power Engineers

Engineering Services

A leader in grid reliability and resiliency, Electric Power Engineers (EPE) has partnered with power and energy clients across the globe for over 50 years. EPE’s client-centric approach delivers unmatched expertise throughout the project lifecycle, with a platform of comprehensive services and proprietary solutions spanning utility engineering, grid and resource integration, reliability and compliance, digitalization, and grid analytics, as well as innovative software partnership with its sister company ENER-I.AI, Inc. Committed to designing and developing the grid of the future, EPE’s highly experienced team of detail-oriented consultants are passionate about helping clients address complex engineering and grid modeling challenges, bridge gaps, and gain visibility into the evolving complexities of the grid. Learn more about EPE at epeconsulting.com

Company details

Company typeScaleup
IndustryEngineering Services
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Overview:

We are designing the grid of the future! 

We are seeking an experienced Senior Site Reliability Engineer (SRE) to join our engineering organization. The ideal candidate will combine strong software engineering and cloud infrastructure expertise to improve the reliability, scalability, security, and operational efficiency of our systems. This role will focus on building resilient AWS and Kubernetes platforms, defining and measuring service reliability, automating operational work, and improving how we detect, respond to, and learn from production incidents. 

The Senior SRE will work closely with software engineering, platform, security, QA, and product teams to establish reliability standards and ensure production services can scale safely as the business grows. 

Responsibilities:

How you can make an impact:

  • Service Reliability & Availability: Define, measure, and improve service reliability using service-level indicators (SLIs), service-level objectives (SLOs), error budgets, availability targets, and capacity planning.
  • Observability: Build and maintain monitoring, logging, tracing, dashboards, and alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, or similar platforms. Ensure alerts are actionable and aligned to customer and service impact.
  • Incident Response: Participate in and improve production incident response, including on-call practices, troubleshooting, escalation, communication, root-cause analysis, and blameless post-incident reviews.
  • Operational Excellence: Improve runbooks, documentation, production readiness reviews, change management, operational standards, and engineering practices that reduce risk and improve system maintainability.
  • Cost & Efficiency: Optimize infrastructure for reliability and performance while maintaining responsible cloud spend and supporting FinOps initiatives.
  • Performance & Capacity: Analyze system performance, resource utilization, latency, throughput, and growth trends; identify bottlenecks and implement scalable solutions before they become production issues.
  • Automation & Toil Reduction: Identify repetitive operational work and replace it with reliable automation using Python, Bash, CI/CD tooling, and platform APIs.
  • CI/CD & Release Reliability: Build and improve deployment pipelines that support safe, repeatable releases through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.
  • Security & Compliance: Apply cloud security and operational controls including least-privilege IAM, encryption, network security, secrets management, patching, auditability, and compliance requirements.
  • Infrastructure as Code: Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
  • Kubernetes Reliability: Operate and improve Kubernetes-based platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, and workload resilience.
  • Cloud Platform Engineering: Architect, operate, and optimize AWS services such as EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related services.
  • Resilience & Disaster Recovery: Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities. Conduct reliability testing and failure exercises where appropriate.
  • Cross-Functional Collaboration: Partner with development teams to improve application operability, reliability, instrumentation, deployment patterns, and production readiness.
Qualifications:

Bring your passion, here's what’s needed:

Required Skills and Qualifications 

  • Experience: 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, including significant production ownership.
  • AWS: Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
  • Terraform: Advanced proficiency with Terraform, including reusable modules, remote state, dependency management, environment design, code review, and infrastructure lifecycle management.
  • Observability: Experience with metrics, logs, traces, dashboards, alerting, and production telemetry using platforms such as Prometheus, Grafana, CloudWatch, New Relic, ELK/OpenSearch, or similar tools.
  • Reliability Engineering: Practical knowledge of SRE concepts such as SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and reducing operational toil.
  • Kubernetes: Deep understanding of Kubernetes architecture and operations, including EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and troubleshooting.
  • Linux & Systems: Strong Linux systems knowledge and the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
  • Programming & Automation: Proficiency in Python, Bash, Go, or another general-purpose language used to build operational tooling and automation.
  • Incident Management: Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective-action tracking.
  • CI/CD: Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable tooling.
  • Security: Working knowledge of cloud security best practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
  • Communication: Strong written and verbal communication skills, with the ability to collaborate effectively across engineering and business teams.

Preferred Qualifications 

  • AWS certifications such as AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer - Professional.
  • Experience with GitOps practices and tools such as Argo CD or Flux.
  • Experience designing or participating in formal on-call rotations and incident management programs.
  • Familiarity with chaos engineering, resilience testing, or game-day exercises.
  • Experience with service meshes, distributed systems, and microservice architectures.
  • Knowledge of database operations for technologies such as Amazon RDS, DynamoDB, and PostgreSQL.
  • Experience with multi-account AWS environments, landing zones, governance, or large-scale cloud platform design.
  • Experience with infrastructure cost optimization, cloud financial management, or FinOps practices.
  • Familiarity with security and compliance frameworks such as SOC 2, ISO 27001, PCI DSS, or similar standards.

What Success Looks Like 

  • Production services become more measurable, reliable, and resilient over time.
  • Operational toil and recurring incidents are reduced through engineering and automation.
  • Teams have clear SLOs, useful dashboards, actionable alerts, and well-understood operational ownership.
  • Infrastructure and deployment changes are repeatable, observable, secure, and low risk.
  • Incidents produce meaningful learning and durable improvements rather than recurring fixes.
  • Cloud capacity and cost are proactively managed without compromising reliability.

Why Join Us? 

  • Work on modern cloud and reliability engineering challenges in a fast-paced, innovative environment.
  • Help shape reliability standards and engineering practices for mission-critical systems.
  • Collaborate with talented engineers across software, cloud, security, and product disciplines.
  • Competitive salary, comprehensive benefits, and opportunities for professional growth.
  • Flexible remote or hybrid work options.

 

Be a part of an innovative team shaping the grid of the future through advanced energy intelligence.  For more than half a century, Electric Power Engineers (EPE) has partnered with power and energy clients across the globe, providing consulting expertise and energy intelligence software solutions for complex engineering and grid modeling challenges. As leaders in the renewables space, we are focused on building a modern, secure, and resilient grid.   Join us in making an impact on the communities we serve and the environment in which we live. Together we can transform the future of energy.  

 

How we support you:

  • Comprehensive health and wellness benefits including medical, dental, and vision with 100% premium coverage for you
  • Generous PTO and paid holidays
  • MyShare Employee Ownership Program
  • Work with industry leaders
  • 401K, up to a 4% match (100% vested from day 1)

 

 

Location: This position will be located in City, State

Travel:  Occasional travel may be needed (10% or less)

 

EPE is an equal opportunity/AA/Disability/Veteran employer. The EEO is the Law poster, and its supplement are available using the following links: EEOC is the Law Poster

 

 

Third-Party Recruiting Notification

EPE does not accept unsolicited resumes from third-party recruiters. Any unsolicited third-party resumes forwarded by recruiters to EPE via our career page or to any of our managers or employees will be considered public information, may be treated as a direct application from the person identified in the resume, and will not be eligible for placement fee payment to the agency. EPE will not pay a fee to a third-party recruiter or agency without a previously signed third-party agreement and has not coordinated their recruiting activity with the appropriate member of the Talent Acquisition team. 

 

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Electric Power Engineers

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.