Logo for Mirantis

Senior Site Reliability Engineer (Golang / Kubernetes)

Role overview

Qualifications

  • 5+ years in SRE, platform reliability, or a closely related software/infrastructure role.
  • Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs or services in production.
  • Demonstrated experience defining SLIs/SLOs and error budgets for real production systems.
  • Hands-on experience with observability tooling β€” metrics, logging, and tracing (e.g., Prometheus/VictoriaMetrics, OpenTelemetry, Grafana).

Responsibilities

  • Define SLIs and SLOs based on the signals available across the platform β€” Kubernetes, bare-metal hosts, and NVIDIA infrastructure.
  • Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems.
  • Establish alerting and error-budget practices that maximize signal and minimize noise.
  • Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected.

Key facts

  • Remote from: United States
  • Full time
  • Senior (5-10 years)
  • Site Reliability Engineer (SRE)
  • English

Hard skills

Other skills

  • Communication

About the company

Mirantis logo

Mirantis

Cloud Computing & Infrastructure (IaaS/PaaS)

Mirantis helps organizations ship code faster on public and private clouds. The company provides a public cloud experience on any infrastructure to the data center to the edge. With Lens and Docker Enterprise Container Cloud, Mirantis empowers a new breed of Kubernetes developers by removing infrastructure and operations complexity and providing one cohesive cloud experience for complete app and devops portability, a single pane of glass, and automated full-stack lifecycle management with continuous updates. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, Nationwide Insurance, PayPal, Reliance Jio, Splunk, and STC. Learn more at www.mirantis.com.

Company details

Company typeSME
IndustryCloud Computing & Infrastructure (IaaS/PaaS)
Company size501 - 1000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environmentβ€”on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Define what reliability means for a GPU-accelerated AI platform and make it measurable. You will own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) β€” deriving meaningful SLIs from the signals the platform already emits, and exposing them to Platform Administrators through a clean API. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.

About the Role

We are looking for a Senior SRE who thinks past dashboards to the contract between a platform and its operators. The right candidate can look at raw telemetry from Kubernetes, bare metal, and NVIDIA infrastructure, decide which signals actually predict user-visible reliability, and turn them into SLIs and SLOs that operators can act on. You are equally comfortable writing the service that exposes those SLIs through an API and reasoning about error budgets, alerting quality, and signal-to-noise. You should be self-directed, able to own reliability definitions end to end, and effectively communicate them across teams.

Responsibilities

  • Define SLIs and SLOs based on the signals available across the platform β€” Kubernetes, bare-metal hosts, and NVIDIA infrastructure (BMC, InfiniBand, NVLink, UFM).

  • Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems.

  • Establish alerting and error-budget practices that maximize signal and minimize noise.

  • Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected.

  • Diagnose reliability and performance issues across the observability stack and drive their resolution.

Qualifications

Required Qualifications

  • 5+ years in SRE, platform reliability, or a closely related software/infrastructure role.

  • Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs or services in production.

  • Demonstrated experience defining SLIs/SLOs and error budgets for real production systems.

  • Hands-on experience with observability tooling β€” metrics, logging, and tracing (e.g., Prometheus/VictoriaMetrics, OpenTelemetry, Grafana).

  • Solid understanding of Kubernetes and the signals it and its workloads emit.

  • Strong written and verbal communication with technical audiences.

Preferred

  • Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM).

  • Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API.

  • Familiarity with VictoriaMetrics/VictoriaLogs at scale.

  • Proven experience in sovereign or high-security air-gapped environments.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Mirantis

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.