Logo for Xideral

Senior AWS Site Reliability Engineers - 2850

Role overview

Qualifications

  • Minimum of 5 years of experience as a Senior Site Reliability Engineer
  • Bachelor’s degree or higher
  • Fluent in English (Advanced)
  • Excellent communication, empathy, commitment, leadership

Responsibilities

  • Own and improve the reliability of cloud-based services and supporting infrastructure
  • Lead incident response activities, including triage, escalation, mitigation, and service restoration
  • Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis
  • Build automation to reduce manual effort and improve operational efficiency

Key facts

Hard skills

Other skills

  • Communication
  • Empathy
  • Leadership
  • Teamwork

About the company

Xideral logo

Xideral

IT Services & IT Consulting

We are a proudly Mexican corporation, with 20 years of experience and operations in Mexico, LATAM and Canada. We have a highly qualified team that allows us to create tailored solutions, implementations and provide services in the IT area, with the aim of achieving the goals of our customers and support their growth. We are committed to the educational and economic development of our country, so we design employment de-centralization programs that help reduce the migration of young people from disadvantaged regions, which creates a chain of welfare for the communities.

Company details

IndustryIT Services & IT Consulting
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Seeking Experienced Senior AWS Site Reliability Engineers for Exciting Projects – Remote in Mexico

We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms, with a focus on production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.

Key Responsibilities:

  • Own and improve the reliability of cloud-based services and supporting infrastructure.
  • Participate in on-call rotations and support production systems outside normal business hours.
  • Lead incident response activities, including triage, escalation, mitigation, and service restoration.
  • Drive blameless postmortems and ensure corrective actions are tracked to closure.
  • Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
  • Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
  • Support and improve cloud and container platforms across AWS and Azure.
  • Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
  • Build automation to reduce manual effort and improve operational efficiency.
  • Configure and improve monitoring, alerting, logging, diagnostics, and observability.

Technical Skills Required:

With over 5 years of experience as a Senior Site Reliability Engineer, you must be proficient in the following technical skills:

  • Strong hands-on experience with AWS and Azure cloud platforms.
  • Strong experience with Terraform for Infrastructure as Code (IaC).
  • Experience with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
  • Strong hands-on experience with Docker and Kubernetes.
  • Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
  • Strong production support experience, including incident management, Root Cause Analysis (RCA), postmortems, and runbook creation.
  • Strong observability experience, including monitoring, alerting, logging, diagnostics, and performance analysis.
  • Good understanding of cloud networking, security, access controls, and InfoSec practices.
  • Experience with version control, branching, merging, pull requests, and conflict resolution.
  • Understanding of cloud cost optimization and resource utilization.

Good-to-Have Skills:

  • Experience with microservice-based platforms.
  • Experience with Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar tools.
  • Scripting or programming experience using Python, Bash, Go, or Java.
  • Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
  • Experience with disaster recovery testing and production readiness reviews.
  • Prior experience mentoring junior engineers or leading technical troubleshooting.

Qualifications:

  • Bachelor’s degree or higher.
  • Fluent in English (Advanced).
  • Excellent communication, empathy, commitment, leadership, teamwork, and a proactive attitude.

Location & Schedule:

  • Remote work from Mexico.
  • Preferred hybrid model in Guadalajara, Jalisco, with expected onsite attendance 2 days per week.
  • Work hours Monday to Friday, 09:00 – 18:00.
  • Advanced English skills are mandatory, and only residents of Mexico.

Benefits:

  • Attractive Salary + Premium Benefits
  • Performance bonuses,  grocery coupons, and savings are found.
  • Aguinaldo,  premium vacations,  and vacations paid
  • SGMM Medical insurance, family, and  Life insurance.

Candidates must include their compensation expectations in their applications and resumes in English.

Interested? Apply now through this link:

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Xideral

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.