Logo for Nebius Group

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Role overview

Qualifications

  • Deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and infrastructure-as-code practices
  • Scripting proficiency in Python or Bash and experience designing alerts and SLOs for high-throughput APIs
  • Production experience with distributed back-ends and GPU-heavy workloads (e.g., vLLM, Triton, Ray) or accelerator stacks
  • Background in MLOps or model-hosting platforms and a focus on building self-healing systems and reliability engineering

Responsibilities

  • Own the reliability, performance, and observability of the entire inference stack by designing and refining telemetry pipelines (metrics, logs, traces) to yield actionable insights
  • Tune Kubernetes autoscalers and craft Terraform modules to maximize GPU utilization and cluster resilience
  • Harden request-routing and retry logic so transient failures are invisible to users and contribute to self-healing behavior
  • Respond to incidents with automation and runbooks to detect, isolate, and remediate problems quickly and lead post-mortem analysis to prevent recurrence

About the company

Nebius Group logo

Nebius Group

Cloud Computing & Infrastructure (IaaS/PaaS)

Nebius Group is building one of the largest commercially available AI infrastructure businesses, based in Europe. Our core business is an AI-centric cloud platform built for intensive AI workloads. We are building full-stack infrastructure to service the explosive growth of the global AI industry, including large-scale GPU clusters, cloud platforms and tools and services for developers. Our headquarters and main R&D presence is in Amsterdam, with additional R&D and commercial hubs in Europe, North America and Israel. Our businesses operate and serve customers on four continents. Our name reflects our mission and values as a company. Combining nebula, the Latin word for cloud, with a reference to the endless Möbius strip, the name Nebius embodies our belief in the endless possibilities offered by cloud technologies. In addition to our core AI infrastructure business, we are developing three other businesses that operate under their own distinctive individual brands: • Toloka AI – a data partner for all stages of AI development from training to evaluation; • TripleTen – a leading edtech player in the US and certain other markets, re-skilling people for careers in tech; • Avride – one of the most experienced teams developing autonomous driving technology for self-driving cars and delivery robots. The group is led by Arkady Volozh, the entrepreneur and visionary co-founder of Yandex.

Company details

IndustryCloud Computing & Infrastructure (IaaS/PaaS)
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Why work at Nebius
Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field.

Where we work
Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 1400 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team.

Token Factory is a part of  Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that promise, we need an engineer who can make the platform behave flawlessly under extreme load and recover gracefully when the unexpected happens.

In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets.

Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about.

If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story.

What we offer 

  • Competitive salary and comprehensive benefits package.
  • Opportunities for professional growth within Nebius.
  • Flexible working arrangements.
  • A dynamic and collaborative work environment that values initiative and innovation.

We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Nebius Group

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.