Logo for The Home Depot

Staff Software Engineer - Customer Reliability Engineering (Remote)

Role overview

Qualifications

  • 8+ years of relevant professional experience in Cloud Operations, Site Reliability Engineering, DevOps, or Software Engineering
  • Deep expertise designing, deploying, and operating high-availability, multi-region production architectures
  • Strong software engineering background with proficiency in Go, Python, or Java
  • Mastery of Infrastructure as Code and CI/CD pipeline automation

Responsibilities

  • Develops, tests, deploys, and maintains software to ensure reliability and efficiency
  • Creates new and better ways for the organization to be successful, ensuring user stories are developer ready
  • Fields questions from product and engineering teams, helping grow junior engineers
  • Defines Service Level Objectives and drives blameless post-incident reviews

Key facts

Hard skills

Other skills

  • Mentorship
  • Communication
  • Collaboration
  • Social Skills
  • Adaptability

About the company

The Home Depot logo

The Home Depot

Retail – Home Improvement & Building Supplies

The Home Depot, the world’s largest home improvement specialty retailer, values and rewards dedicated, knowledgeable, and experienced professionals. We operate more than 2,300 retail stores in all 50 states, the District of Columbia, Puerto Rico, the U.S. Virgin Islands, Guam, Canada, and Mexico. All of our associates have one thing in mind — helping our customers build and improve their homes. Join The Home Depot team today and see for yourself why we are consistently ranked as a top Fortune 500 company.

Company details

Company typeXLarge
IndustryRetail – Home Improvement & Building Supplies
Company size10001

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

With a career at The Home Depot, you can be yourself and also be part of something bigger.

Position Purpose:

The Customer Reliability Engineering team ensures the continuous performance, security, and resilience of various customer profile and communication services that support the company's e-commerce platform and in-store systems. The Staff Reliability Engineer is responsible for leading the design and implementation of foundational systems that ensure the reliability, scalability, performance, and efficiency of the products our customers and associates love. As a Staff SRE, you will serve as a technical anchor, establishing the blueprints for observability, automation, and cloud infrastructure architecture
across the enterprise. You will drive tool selection, configuration, security, resilience, performance tuning, and production monitoring to systematically reduce operational toil. In this role, you will balance rapid product delivery with long-term system stability by defining Service Level Objectives (SLOs), error budgets, and driving a culture of blameless post-incident reviews. As a core player on the engineering team, you are expected to drive technical decisions across teams without formal authority and actively mentor engineers of all experience levels to elevate the organization's engineering standards.


Key Responsibilities:

  • 50% Delivery and Execution - Develops, tests, deploys, and maintains software, with a clear understanding of the value the software is to provide; Takes a broad view when approaching issues; using a global lens; Consistently achieves results, even under tough circumstances; Develops test suites (functional, destructive, etc) to enable success, rapid deployment of code to production; Takes on new opportunities and tough challenges with a sense of urgency, high energy and enthusiasm; Consistently achieves results, even under tough circumstances
  • 10% Learns and Grows - Actively seeks ways to grow and be challenged using both formal and informal development channels; Learns through successful and failed experiment when tackling new problems
  • 20% Plans and Aligns - Creates new and better ways for the organization to be successful; Delivers multi-mode communications that convey a clear understanding of the unique needs of different audiences; Works the Product Team to ensure user stories are developer ready, easy to understand and testable; Collaborates with other team members in agile processes; Relates openly and comfortably with diverse groups of people; Adapts approach and demeanor in real time to match the shifting demands of different situations
  • 20% Supports and Enables - Fields questions from product and engineering teams; Helps grow junior engineers by providing guidance on modern software development frameworks, and leading technical discussions; Notes gaps on the team and provides suggestions for changes to make the team more productive


Direct Manager/Direct Reports:

  • This position typically reports to Software Engineer Manager or Sr. Manager
  • This position typically has 0 Direct Reports


Travel Requirements:

  • No travel required.


Physical Requirements:

  • Most of the time is spent sitting in a comfortable position and there is frequent opportunity to move about. On rare occasions there may be a need to move or lift light articles.


Working Conditions:

  • Located in a comfortable indoor area. Any unpleasant conditions would be infrequent and not objectionable.


Minimum Qualifications:

  • Must be eighteen years of age or older.
  • Must be legally permitted to work in the United States.


Preferred Qualifications:

  • Experience: 8+ years of relevant professional experience in Cloud Operations, Site Reliability Engineering, DevOps, or Software Engineering in a high-scale, distributed environment.
  • Cloud & Infrastructure: Deep expertise designing, deploying, and operating high-availability, multi-region production architectures on Google Cloud Platform (or AWS/Azure).
  • Software Development: Strong software engineering background with production-level proficiency in Go, Python, or Java to build distributed automation, internal developer platforms (IDPs), and custom reliability tooling.
  • Observability & Monitoring: Deep expertise in building observability solutions using tools such as Datadog, Prometheus, Grafana, or Splunk. Proven ability to define and implement SLIs, SLOs, and alerting strategies.
  • Automation & IaC: Mastery of Infrastructure as Code (e.g., Terraform, CloudFormation) and CI/CD pipeline automation (e.g., GitHub Actions, Jenkins).
  • Containerization: Advanced operational experience with Kubernetes, container orchestration, and microservices architectures.
  • Incident Management: Strong background in leading incident response coordination, conducting root-cause analysis, and driving systemic architectural improvements through blameless post-mortems.
  • Technical Leadership: Demonstrated ability to lead the technical direction of complex, cross-team initiatives, navigating ambiguous challenges and driving scalable solutions from concept to production.
  • Mentorship: Proven track record of mentoring junior and mid-level engineers, fostering technical excellence, and raising the engineering bar through architecture and code reviews.
  • Collaboration: Strong communication skills with the ability to partner effectively across Product, UX, Architecture, Security, and Engineering teams to influence technical direction and prioritize reliability.


Minimum Education:

  • The knowledge, skills and abilities typically acquired through the completion of a bachelor's degree program or equivalent degree in a field of study related to the job.


Preferred Education:

  • No additional education


Minimum Years of Work Experience:

  • 3


Preferred Years of Work Experience:

  • No additional years of experience


Minimum Leadership Experience:

  • None


Preferred Leadership Experience:

  • None


Certifications:

  • None


Competencies:

  • Global Perspective
  • Manages Ambiguity
  • Nimble Learning
  • Self-Development
  • Collaborates
  • Cultivates Innovation
  • Situational Adaptability
  • Communicates Effectively
  • Drives Results
  • Interpersonal Savvy

For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:
 

The pay range for this position is between $90,000.00 - $190,000.00

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Software Engineer Related jobs

Other jobs at The Home Depot

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.