Job Title - Senior Site Reliability Engineer/ Lead
Experience Required- 7 to 9 Years exp
Shift Timing - 7:30 PM IST - 4:30 AM IST
Job Location- Remote
About the Role
We are hiring a Site Reliability Engineer who can do three things well: take a production failure down to its actual root cause and close it : including when the cause turns out to be the alert itself rather than the service; build the automation that stops that class of failure from ever needing a human again; and own a project end-to-end : taking a service onto our release tooling, standing up a new tenant's queues, promoting a set of SLOs : from scope to production.
We are a fundamentally proactive team. We care about release process and tooling because that's how we prevent outages before they happen. We invest in tooling because tooling : not heroics : is how we stop the next incident. This isn't a reactive ops role; it's a role for someone who treats every incident as an unfinished automation task.
What You'll Do
Own incidents to true root cause. Go beyond "the alert fired, service restarted " : dig into logs, metrics, and system state until you find the actual cause, even when that cause is a bad alert, a flawed threshold, or a gap in the release process rather than the service itself
Close the loop with automation. For every recurring failure mode, build the script, tool, or process change that removes the need for a human to intervene the next time it happens
Own projects end-to-end. Take full ownership of scoped initiatives : onboarding a service onto our release tooling, standing up queues for a new tenant, defining and promoting a set of SLOs : from initial scope through to production, without requiring hand-holding
Drive release process and tooling improvements , because we believe strong release discipline is one of the highest-leverage ways to prevent outages
Improve observability and alerting quality : when an alert is noisy, wrong, or unclear, fix the alert, not just the runbook Explore AI-assisted automation : apply LLMs, Claude skills, and RAG-based approaches to triage, diagnostics, and knowledge retrieval, reducing repetitive manual investigation during incidents
Participate in and improve on-call rotations , with an emphasis on reducing on-call burden over time through automation, not just responding faster
Write clear, precise incident writeups : root cause, timeline, and the concrete follow-up (usually automation) that prevents recurrence
Collaborate with service owners and platform teams to scope and execute cross-team projects, and to make sure fixes and automation actually get adopted
Mentor junior engineers on debugging methodology, root-cause analysis, and building durable automation rather than one-off fixes
What We're Looking For Required
7-9 years of experience in SRE, DevOps, or infrastructure/production engineering roles
Strong track record of driving production issues to true root cause : not just mitigating symptoms
Proven ability to build automation (scripts, tools, services) that durably eliminates a class of manual or repetitive work Demonstrated experience owning technical projects independently from scoping through to production delivery Strong Python skills : comfortable building and maintaining automation, tooling, and scripts used by other engineers
Advanced skills reading and correlating logs, metrics, and traces to diagnose issues across multiple interdependent systems Extensive experience with release engineering, deployment tooling, or CI/CD pipelines
Excellent written communication : clear incident writeups, project scoping docs, and technical updates that influence direction
across teams
Experience mentoring engineers and leading technical initiatives
Preferred
Exposure to applying LLMs for operational use cases : building Claude skills/agents, RAG pipelines for internal knowledge/runbook retrieval, or AI-assisted triage and diagnostics
Experience with observability stacks (Prometheus, Grafana, Datadog) and designing/tuning alerts and SLOs Familiarity with queueing systems and multi-tenant infrastructure
Experience with cloud infrastructure (AWS, GCP, or Azure)
Exposure to CI/CD and release tooling systems (Jenkins, GitHub Actions, Spinnaker, or similar) : bonus if you've contributed to internal release tooling
Working knowledge of Kubernetes fundamentals (workload troubleshooting, deployments) : note: this role does not own cluster operations, upgrades, or node group management
Technologies You'll Work With
Python · LLMs / Claude / RAG-based tooling · CI/CD tooling · Release tooling · Observability stacks (Prometheus/Grafana or similar) · Cloud infrastructure (AWS/GCP/Azure) · Queueing systems