Logo for Level AI

Senior Site Reliability Engineer (Remote, India) EST

Role overview

Qualifications

  • 4-5 years of hands-on systems experience
  • Production experience in Python, Go/Rust
  • Fluency in GCP, IaC (Terraform), CI/CD
  • Familiarity with GPU workload understanding

Responsibilities

  • Own infrastructure cost efficiency and FinOps
  • Run a structured experimentation program on on-premise GPU clusters
  • Build tooling, dashboards, and processes for backend teams
  • Take on defined security workstreams

Key facts

Hard skills

Other skills

  • Problem Solving

About the company

Level AI logo

Level AI

Customer Experience & Contact Centers

Our state-of-the-art AI-native solutions are designed to drive efficiency, productivity, scale, and excellence in sales and customer service. With a focus on automation, agent empowerment, customer assistance, and strategic business intelligence, we are dedicated to helping our clients exceed customer expectations and drive profitable business growth. Companies like Affirm, Carta, Vista, Toast, Swiss Re, ezCater, etc. use Level AI to take their business to new heights with less effort.

Company details

Company typeScaleup
IndustryCustomer Experience & Contact Centers
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Please note: This role requires working in the EST Time Zone (usually from 8 PM to 4 AM IST) in order to add technical coverage to US customers.

About Level AI

Level AI is on a mission to turn every customer interaction into a strategic advantage. Our AI-native platform helps enterprises transform contact centers from cost centers into engines of customer intelligence, operational efficiency, and business growth. By combining advanced AI with deep domain understanding of customer experience, Level AI empowers teams to unlock actionable insights, automate workflows, and deliver more consistent, higher-quality support across the customer journey.

Headquartered in Mountain View, California, Level AI is a Series C company backed by leading investors, including Battery Ventures and ENIAC. Our platform leverages Large Language Models and Custom Small Language Models (SLMs) to power AI Agents across the entire CX journey—customer-facing agents, agent-assist, and backend automation—along with deep conversation analytics for QA, coaching, and insights.

About the role

The Senior SRE will be positioned at the intersection of backend engineering, infrastructure operations, and FinOps. The role is explicitly broader than a traditional DevOps engineer and explicitly more hands-on than a pure architect.

What you'll be liable for:

Infrastructure cost efficiency and FinOps. Own the continued reduction of Kubernetes overprovisioning, drive right-sizing programs, and maintain the cost telemetry that backend teams use to make decisions.

GPU throughput optimization. Run a structured experimentation program on on-premise GPU clusters, partnering with AI service owners. Led by the Engineering leadership, with this role providing the experimental bandwidth.

Backend enablement, not ownership absorption. Build the tooling, dashboards, and processes that let backend teams from other groups own their own cost and reliability budgets. The deliverable is leverage, not headcount-shaped work.

Reliability instrumentation. As the infra team owns most of the instrumentation across new and offline flows, this role takes a central seat in making sure that surface area is captured properly for both cost-at-scale and reliability.

Selective security workstreams. Take on a defined slice of the active security work so that senior DevOps engineers are not the single point of execution for security-adjacent platform changes.


We'd love to explore more about you if you have:

This role explicitly requires 4-5 years of hands-on systems experience. We are not looking for someone who will lean entirely on AI tooling to discover what to do; we are looking for someone who already knows what to ask and can use AI tooling as a force multiplier on top of that judgement.

Backend engineering depth: production experience in Python, Go/Rust, comfortable owning services end to end, able to read and reason about backend code across teams.

Kubernetes at scale: scheduler behaviour, resource requests/limits, HPA/VPA, node pool design, cost-aware autoscaling (Cast AI, Karpenter, or equivalent).

Cloud and on-premise infrastructure: GCP fluency, IaC (Terraform), CI/CD, and comfort operating in hybrid setups, including on-prem GPU clusters.

GPU workload understanding: familiarity with throughput profiling, batching, KV-cache behavior, inference server tuning, and GPU utilisation metrics.

Observability and reliability: metrics, traces, logs, SLOs, and the discipline to instrument systems properly rather than reactively.

FinOps mindset: demonstrated history of converting infrastructure choices into measurable cost outcomes.

Security baseline: able to take on platform-security workstreams without requiring constant handoff to the DevOps team.

Can work in the EST time zone (A must)

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Level AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.