Logo for Sur La Table

Site Reliability Engineer

Role overview

Qualifications

  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS
  • Bachelor's degree in computer science or similar, or equivalent experience

Responsibilities

  • Work on service resiliency, performance tuning, and system design across Backcountry's platform
  • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
  • Leverage AI-assisted engineering tools to investigate, automate, and ship fixes across infrastructure and application repositories
  • Monitor system health and capacity, taking proactive action to fix problems before they occur

About the company

Sur La Table logo

Sur La Table

Retail – Furniture & Home Furnishings

Sur La Table was founded in Seattle in 1972 by Shirley Collins, a woman with a passion for food and a fondness for community. Living in Seattle, she fell in love with Pike Place Market with its inspiring blend of products, artisans, and farmers. To her, it was a special gathering place for food lovers and culinary visionaries alike. When Shirley opened her first store in Pike Place Market, she was determined to assemble the best selection of cookware, gadgets, linens and books—even importing exclusive specialty items from France, her favorite culinary destination. Using the market as her inspiration, she thoughtfully filled her store with cooking tools that would bring people together in the kitchen and around the table. This sense of connection and love of French cuisine inspired the name Sur La Table, which simply means “on the table.” Since then we’ve grown to 56 stores across America, with the largest avocational cooking program in the U.S. But some things haven’t changed: We’re still the place for an unsurpassed selection of exclusive and premium-quality goods for the kitchen and table. We’re still passionate about cooking and entertaining, eager to share all we know. Whether the job entails interacting with our customers on a daily basis or providing the vital behind-the-scenes support, we’re all here for the same reason – to create happiness through cooking and sharing good food.

Company details

IndustryRetail – Furniture & Home Furnishings
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Adventure is our Culture. Join a team that celebrates a lifestyle as bold as the terrain we love. At Backcountry, we are rooted in adventure, recognition, and wellbeing on and off the mountain. We spotlight employee stories, celebrate milestones, and offer exclusive outdoor perks. Whether you are at HQ, in a retail store, or remote, you will be part of a team that thrives on energy, exploration, and connection.

 

Reports to: Gustavo Arguedas (Site Reliability Manager)

Location: Remote - Costa Rica


About the Role

Backcountry's online platform serves as the backbone of our customer experience, and this role exists to ensure its reliability, performance, and scalability. As Site Reliability Engineer, you will partner with software engineering, DevOps, and IT operations teams to optimize systems and applications across a multi-cloud stack.

Within 6–12 months, you will have contributed meaningful improvements to service resiliency and observability, reduced operational toil through automation, and established yourself as a trusted partner to development and infrastructure teams.

This is a lean team. You will own a lot, move fast, and make decisions with full end-to-end responsibility.


What You'll Do
  • Work on service resiliency, performance tuning, and system design across Backcountry's platform
  • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
  • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories
  • Reduce toil by designing and implementing automation
  • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads
  • Monitor system health and capacity, taking proactive action to fix problems before they occur
  • Collaborate with engineering teams to build, deploy, and support features
  • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services
  • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy
  • Participate in the on-call support rotation within the SRE team

  • Required Qualifications
  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written

  • Preferred Qualifications
  • Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements
  • Familiarity with PCI-scoped or other regulated environments
  • Previous experience working in ecommerce environments
  • Professional certifications: GCP, CKA, or AWS

  • Why Join

    The people who do best here are builders. They take ownership, move fast, and want to see the direct impact of their work.

  • Cross-Functional Impact: Your work directly affects platform reliability for every customer and every team that depends on Backcountry's systems.
  • Modern Tech Stack: Work across a multi-cloud environment (GCP and AWS) with modern observability tooling, AI-assisted engineering, and GitOps workflows.
  • End-to-End Ownership: Own projects from investigation through implementation — you will ship automation, improve resiliency, and see the results in production.
  • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

  • Interview Process
    1. Recruiter Screen - A 30-minute conversation with our recruiting team to align on the role, your background, and what you are looking for.
    2. Hiring Manager Interview - Conversation with the Site Reliability Manager focused on your SRE experience, approach to incident management, and team fit.
    3. Technical/Case Discussion - A deeper dive into infrastructure, observability, and problem-solving scenarios relevant to the role.
    4. Reference Checks - Conducted in parallel with the final stages where possible.
    5. Offer - We move quickly for the right candidate.

    Interview process is subject to change. Any updates will be communicated promptly and clearly.

    CSC Generation is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected by law.

    The CSC Generation family of brands is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or accommodation due to a disability, please contact hrbenefits@cscshared.com.

    Apply once. Then go straight to the hiring manager.

    After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

    MR

    Marcus Rivera

    Chief Revenue Officer

    m.rivera@company.com
    linkedin.com/in/marcusrivera
    Unlocked after you apply
    ·

    Site Reliability Engineer (SRE) Related jobs

    Other jobs at Sur La Table

    Premium

    Reach out to the hiring manager directly.

    Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

    • Full match report with fit score and gaps
    • Career diagnostics on how recruiters read you
    • Curated company matches and warm intros
    • 48h early access to new roles

    Cancel anytime.