Logo for STN Incorporated

Site Reliability Engineer

Role overview

Qualifications

  • 5+ years in SRE, DevOps, or production engineering roles
  • Strong programming skills in Go, Python, or both
  • Hands-on experience operating Kubernetes-based platforms at scale
  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

Responsibilities

  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs
  • Build and maintain the observability stack including metrics, logs, traces, and alerting
  • Lead incident response and chair post-incident reviews
  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

Key facts

  • Remote from: California (USA)
  • Full time
  • Senior (5-10 years)
  • Site Reliability Engineer (SRE)
  • English

Hard skills

Other skills

  • Mentorship
  • Collaboration

About the company

STN Incorporated logo

STN Incorporated

IT Services & IT Consulting

At STN, we don’t just deliver technology, we build the foundation that modern organizations run on. From enterprise IT to AI infrastructure, we design, operate, and support systems that are reliable, secure, and built for performance. But what sets us apart isn’t just our stack it’s how we show up. We don’t believe in one-size-fits-all. We don’t drop in hardware and disappear. We work side by side with our customers to understand what they actually need and build solutions that fit, flex, and scale as they grow.Whether you're running business-critical systems, deploying AI models, or training large-scale workloads on NVIDIA GPUs, we’re here to make sure your infrastructure isn’t just running, it’s working for you. Our team brings deep technical expertise, hands-on support, and a people-first mindset to everything we do. Because we believe technology should unlock potential, not create more problems. STN exists to help teams thrive in complex environments with custom engineering, real partnership, and a clear plan forward.

Company details

IndustryIT Services & IT Consulting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities

  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications

  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications

  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at STN Incorporated

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.