Logo for Nagarro

Senior Site Reliability Engineer (AWS Cloud)

Role overview

Qualifications

  • 8+ years of experience working within the cloud environment, in roles such as SRE or Cloud Platform/Reliability Engineer
  • Strong experience in cloud development and multi-cloud environments, preferably with a strong exposure to AWS cloud
  • Knowledge of cloud architecture, scalability, and high-availability design
  • Hands-on experience with Kubernetes and container orchestration

Responsibilities

  • Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks
  • Architect and implement end-to-end automation pipelines to eliminate manual intervention
  • Define, monitor, and improve critical system health indicators (SLIs/SLOs)
  • Lead collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC)

About the company

Nagarro logo

Nagarro

IT Services & IT Consulting

Nagarro helps future-proof your business through a forward-thinking, fluidic, and CARING mindset. We excel at digital engineering and help our clients become human-centric, digital-first organizations, augmenting their ability to be responsive, efficient, intimate, creative, and sustainable. Today, we are 18,300 experts across 37 countries, forming a Nation of Nagarrians, ready to help our customers succeed.

Company details

Company typeXLarge
IndustryIT Services & IT Consulting
Company size10001

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Company Description

👋🏼 We're Nagarro.

We are a digital product engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18 000+ experts across 39 countries, to be exact). Our work culture is dynamic and non-hierarchical. We're looking for great new colleagues. That's where you come in!

By this point in your career, it is not just about the tech you know or how well you can code. It is about what more you want to do with that knowledge. Can you help your teammates proceed in the right direction? Can you tackle the challenges our clients face while always looking to take our solutions one step further to succeed at an even higher level? Yes? You may be ready to join us.

Job Description

  • Architect & Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks.
  • Architect and implement end-to-end automation pipelines to eliminate manual intervention, actively identifying and reducing technical toil.
  • Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage, making data-driven architectural recommendations.
  • Lead Collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC).
  • Own Incident Response Management: Design robust detection mechanisms, triage critical incidents, lead deep Root Cause Analysis (RCA), and implement long-term preventative engineering solutions.
  • Simplify Complex Systems: Continually audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity.

Qualifications

  • 8+ years of experience working within the cloud environment, in roles such as SRE (Site Reliability Engineer) or Cloud Platform/Reliability Engineer
  • Strong experience in cloud development and multi-cloud environments, preferably with a strong exposure to AWS cloud
  • Knowledge of cloud architecture, scalability, and high-availability design
  • Hands-on experience with Kubernetes and container orchestration
  • Experience with Terraform and Infrastructure as Code (IaC)
  • Experience designing automation and CI/CD pipelines to reduce operational toil
  • Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
  • Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
  • Ability to identify performance bottlenecks, single points of failure, and architectural risks

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Nagarro

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.