Logo for SWCA Environmental Consultants

Senior Site Reliability Engineer

Role overview

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or similar role
  • Experience with cloud platforms (AWS, Azure, GCP) and container orchestration tools (Docker, Kubernetes)
  • Proficiency with monitoring and logging tools (Prometheus, Grafana, ELK stack, Datadog)
  • Strong scripting skills in Python, Bash, or Go

Responsibilities

  • Monitor and maintain the health, availability, and performance of critical transportation services
  • Define and track service-level objectives (SLOs), error budgets, and reliability metrics for revenue-critical services
  • Lead incident response, perform root cause analysis, and coordinate resolution across teams
  • Develop and implement automation scripts to streamline operational tasks and improve efficiency

About the company

SWCA Environmental Consultants logo

SWCA Environmental Consultants

Environmental Services

SWCA is a 100% employee-owned environmental firm that offers comprehensive environmental planning, regulatory compliance, and natural and cultural resources management services. Our expanding team of professionals combines scientific expertise with in-depth knowledge of permitting and compliance protocols to achieve technically sound, cost-effective solutions for a full spectrum of environmental projects throughout the U.S. and its territories.

Company details

Company typeLarge
IndustryEnvironmental Services
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Overview:

Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that processes roadway transactions and payments around the clock. As a Senior SRE, you will ensure our systems are highly available, resilient, and scalable, take a leading role in incident response and reliability engineering practices, and help optimize operations across infrastructure and applications.

Responsibilities:
  • System Reliability: Monitor and maintain the health, availability, and performance of critical transportation services.
  • Reliability Standards: Define and track service-level objectives (SLOs), error budgets, and reliability metrics for revenue-critical services. 
  • Incident Management: Lead incident response — perform root cause analysis, coordinate resolution across teams, and drive blameless post-incident reviews and follow-up actions.
  • Automation: Develop and implement automation scripts to streamline operational tasks and improve efficiency.
  • Monitoring & Performance: Set up and maintain monitoring, logging, tracing, and alerting tools (e.g., Prometheus, Grafana, OpenTelemetry) to track service health, performance, and resource utilization.
  • Capacity Planning: Help assess and plan for capacity, scaling infrastructure to meet growing transaction volumes — including stateful systems such as distributed SQL databases and event-streaming clusters (e.g., NATS, Kafka).
  • Collaboration: Work with software engineering, infrastructure, and operations teams to improve the reliability of systems and services.
  • System Optimization: Identify performance bottlenecks, troubleshoot issues, and work on optimizations at both infrastructure and application layers.
  • Continuous Improvement: Contribute to the ongoing improvement of operational processes, documentation, and best practices in the SRE team.
  • Disaster Recovery: Participate in designing and testing disaster recovery plans to ensure the continuity of critical services.

This list of responsibilities might not cover everything you'll end up doing.  

Qualifications:
  • Experience: 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role, preferably in a mission-critical or large-scale environment, including experience leading incident response and mentoring other engineers.
  • Technical Skills:
    • Experience with cloud platforms (AWS, Azure, GCP) and container orchestration tools (Docker, Kubernetes).
    • Proficiency with monitoring and logging tools (Prometheus, Grafana, ELK stack, Datadog, etc.).
    • Strong scripting skills in Python, Bash, or Go.
    • Solid understanding of Linux and Windows administration.
  • Database Knowledge: Familiarity with relational databases (MySQL, PostgreSQL, etc.) and distributed systems.
  • Collaboration & Communication: Excellent teamwork and communication skills, with the ability to work across teams to improve service reliability.
  • Problem-Solving: Strong troubleshooting skills with a proactive, solution-oriented mindset.
  • Experience in Intelligent Transportation: While not required, familiarity with transportation systems, autonomous vehicles, or real-time data systems is a plus.

Preferred Qualifications:

  • Experience with traffic management systems, sensor data processing, or other intelligent transportation systems.
  • Knowledge of infrastructure-as-code tools (e.g., Terraform, Ansible, Helm) and GitOps workflows (e.g., Argo CD).
  • Exposure to CI/CD pipelines, including pipeline-as-code (e.g., Dagger, GitHub Actions), and Git-based version control.
Benefits:

We offer a Total Rewards plan designed with you and your family’s health and wellness in mind that includes: 

  • Paid days off (i.e. vacation, sick days, bereavement leave) 
  • Health and Dental plans 
  • Retirement plans 
  • Employee and Family Assistance Program (EFAP) 
  • Employee referral program 

 

We welcome applicants from all backgrounds, regardless of race, color, religion, sex, veteran status, sexual orientation, gender identity, national origin, age, or disability or any other protected characteristics in accordance with applicable federal, state/provincial, and local laws. We're committed to creating a workplace where everyone feels valued and respected.  

 

We appreciate all responses and will acknowledge only those being considered for an interview. 

We respectfully request no calls or unsolicited resumes from Agencies.   

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at SWCA Environmental Consultants

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.