Logo for NVIDIA

Senior Site Reliability Engineer

Role overview

Qualifications

  • BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field
  • 5+ years of experience as a Site Reliability Engineer (SRE) or similar role
  • Strong understanding of containerization, microservices architecture, and Kubernetes
  • Hands-on experience developing automation using Python, Go, Bash, or similar scripting/programming languages

Responsibilities

  • Monitor, support, and maintain the reliability, availability, and performance of GeForce NOW production services
  • Participate in production incident triage, troubleshooting, and resolution of complex issues
  • Collaborate with software engineering, platform, networking, and infrastructure teams to improve service resilience
  • Design and develop custom tools, automation, and self-service solutions to enhance GeForce NOW platform

About the company

NVIDIA logo

NVIDIA

Semiconductors

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Company details

Company typeXLarge
IndustrySemiconductors
Company size10001

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are now looking for a Sr. Site Reliability Engineer (SRE)! NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s motivated by outstanding technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. NVIDIA is at the forefront of generative AI models, from language to images. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work.

NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its cloud service team for supporting, triaging, and building Geforcenow cloud gaming platform. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. We live SRE practices that are key to product quality, such as limiting time spent on reactive operational work, blameless postmortems, proactive identification of potential outages, and iterative improvements, which all make for interesting and dynamic day-to-day work. The person in this position will be responsible for Service Response and workflow and will drive tools/service development to maintain and improve service SLOs. We partner with Service Owners to drive the reliability of the service. Continuously evaluate operational processes, identify opportunities for improvement, and build custom tools, automation, and self-service solutions that improve the overall reliability and operational efficiency of GeForce NOW.

What You Will Be Doing

  • Monitor, support, and maintain the reliability, availability, and performance of large-scale GeForce NOW production services running across cloud and datacenter environments.

  • Participate in production incident triage, troubleshooting, and resolution of complex infrastructure and application issues. Take part in the team's on-call rotation, including occasional weekend coverage, to ensure timely restoration of customer-facing services.

  • Monitor service health using metrics, logs, traces, and dashboards, and proactively identify reliability, performance, and capacity issues before they impact customers.

  • Collaborate with software engineering, platform, networking, and infrastructure teams to improve operational readiness, reliability, and service resilience.

  • Drive observability initiatives by improving monitoring, alerting, dashboards, and telemetry to enable faster detection and diagnosis of production issues.

  • Scale services sustainably by building automation, eliminating operational toil, and continuously improving deployment, recovery, and operational workflows.

  • Lead and participate in incident response, root cause analysis, and blameless postmortems, driving corrective and preventive actions to improve long-term service reliability.

  • Design and develop custom tools, automation, and self-service solutions that simplify operations, improve engineer productivity, and enhance the overall GeForce NOW platform.

  • Continuously evaluate existing operational processes and identify opportunities to improve service reliability, operational efficiency, and customer experience through engineering-driven solutions.

  • Contribute to the design, deployment, and operation of Kubernetes-based services, ensuring they meet scalability, reliability, and performance requirements.

What we need to see:

  • BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field (or equivalent experience).

  • 5+ years of experience supporting and operating mission-critical production services in a live-site environment as a Site Reliability Engineer (SRE), Production Engineer, or similar role.

  • Strong understanding of containerization, microservices architecture, and Kubernetes, including Kubernetes ecosystem components and operational best practices.

  • Demonstrated ability to troubleshoot complex production issues, identify root causes, and drive issues to resolution.

  • Strong understanding of distributed systems and how complex production environments interact across applications, infrastructure, networking, and cloud services.

  • Experience supporting production operations, including incident management, change management, postmortem reviews, and operational excellence initiatives.

  • Hands-on experience developing automation using Python, Go, Bash, or similar scripting/programming languages.

  • Strong understanding of SLOs, SLIs, error budgets, KPIs, and service reliability best practices.

  • Experience with observability platforms such as Prometheus, Grafana, ELK/OpenSearch, and modern monitoring and alerting solutions.

  • Experience operating services in public cloud environments such as AWS, Azure, GCP, or equivalent cloud platforms.

Ways to stand out from the crowd:

  • Experience supporting large-scale customer-facing cloud or gaming services.

  • Strong Kubernetes operational and troubleshooting expertise.

  • Experience with observability platforms, including Prometheus, Grafana, ELK/OpenSearch, and OpenTelemetry.

  • Strong scripting or programming skills in Python, Go, or similar languages with a focus on automation.

  • Experience driving production incident response, postmortems, and operational excellence initiatives.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at NVIDIA

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.