Logo for Scalence L.L.C.

Senior Site Reliability Engineer/ Lead

Role overview

Qualifications

  • 7-9 years of experience in SRE, DevOps, or infrastructure/production engineering roles
  • Strong track record of driving production issues to true root cause
  • Proven ability to build automation that durably eliminates manual or repetitive work
  • Strong Python skills

Responsibilities

  • Own incidents to true root cause by analyzing logs, metrics, and system state
  • Build automation for recurring failure modes to reduce human intervention
  • Take full ownership of scoped initiatives from initial scope through to production
  • Drive release process and tooling improvements to prevent outages

Key facts

Hard skills

Other skills

  • Communication
  • Mentorship

About the company

Scalence L.L.C. logo

Scalence L.L.C.

IT Services & IT Consulting

In today’s dynamic and competitive market, success hinges on mastering three key areas: Data Intelligence, Business Resilience, and Digital Experience. At Scalence, a global women-owned IT and BPO solutions provider, we specialize in these critical disciplines through our IT Project and Managed Solutions, propelling your business to new heights. Our commitment to Customer Success and Delivery Excellence drives everything we do. Let us help you unlock new opportunities and propel your business forward. Explore how we can make a difference today!

Company details

IndustryIT Services & IT Consulting
Company size501 - 1000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Job Title - Senior Site Reliability Engineer/ Lead
Experience Required- 7 to 9 Years exp
Shift Timing - 7:30 PM IST - 4:30 AM IST
Job Location- Remote

About the Role

We are hiring a Site Reliability Engineer who can do three things well: take a production failure down to its actual root cause and close it : including when the cause turns out to be the alert itself rather than the service; build the automation that stops that class of failure from ever needing a human again; and own a project end-to-end : taking a service onto our release tooling, standing up a new tenant's queues, promoting a set of SLOs : from scope to production.

We are a fundamentally proactive team. We care about release process and tooling because that's how we prevent outages before they happen. We invest in tooling because tooling : not heroics : is how we stop the next incident. This isn't a reactive ops role; it's a role for someone who treats every incident as an unfinished automation task.

What You'll Do

Own incidents to true root cause. Go beyond "the alert fired, service restarted " : dig into logs, metrics, and system state until you find the actual cause, even when that cause is a bad alert, a flawed threshold, or a gap in the release process rather than the service itself
Close the loop with automation. For every recurring failure mode, build the script, tool, or process change that removes the need for a human to intervene the next time it happens
Own projects end-to-end. Take full ownership of scoped initiatives : onboarding a service onto our release tooling, standing up queues for a new tenant, defining and promoting a set of SLOs : from initial scope through to production, without requiring hand-holding
Drive release process and tooling improvements , because we believe strong release discipline is one of the highest-leverage ways to prevent outages
Improve observability and alerting quality : when an alert is noisy, wrong, or unclear, fix the alert, not just the runbook Explore AI-assisted automation : apply LLMs, Claude skills, and RAG-based approaches to triage, diagnostics, and knowledge retrieval, reducing repetitive manual investigation during incidents
Participate in and improve on-call rotations , with an emphasis on reducing on-call burden over time through automation, not just responding faster
Write clear, precise incident writeups : root cause, timeline, and the concrete follow-up (usually automation) that prevents recurrence
Collaborate with service owners and platform teams to scope and execute cross-team projects, and to make sure fixes and automation actually get adopted
Mentor junior engineers on debugging methodology, root-cause analysis, and building durable automation rather than one-off fixes

What We're Looking For Required

7-9 years of experience in SRE, DevOps, or infrastructure/production engineering roles
Strong track record of driving production issues to true root cause : not just mitigating symptoms
Proven ability to build automation (scripts, tools, services) that durably eliminates a class of manual or repetitive work Demonstrated experience owning technical projects independently from scoping through to production delivery Strong Python skills : comfortable building and maintaining automation, tooling, and scripts used by other engineers
Advanced skills reading and correlating logs, metrics, and traces to diagnose issues across multiple interdependent systems Extensive experience with release engineering, deployment tooling, or CI/CD pipelines
Excellent written communication : clear incident writeups, project scoping docs, and technical updates that influence direction
across teams
Experience mentoring engineers and leading technical initiatives

Preferred

Exposure to applying LLMs for operational use cases : building Claude skills/agents, RAG pipelines for internal knowledge/runbook retrieval, or AI-assisted triage and diagnostics
Experience with observability stacks (Prometheus, Grafana, Datadog) and designing/tuning alerts and SLOs Familiarity with queueing systems and multi-tenant infrastructure
Experience with cloud infrastructure (AWS, GCP, or Azure)
Exposure to CI/CD and release tooling systems (Jenkins, GitHub Actions, Spinnaker, or similar) : bonus if you've contributed to internal release tooling
Working knowledge of Kubernetes fundamentals (workload troubleshooting, deployments) : note: this role does not own cluster operations, upgrades, or node group management

Technologies You'll Work With

Python · LLMs / Claude / RAG-based tooling · CI/CD tooling · Release tooling · Observability stacks (Prometheus/Grafana or similar) · Cloud infrastructure (AWS/GCP/Azure) · Queueing systems

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Scalence L.L.C.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.