Logo for Xideral

Site Reliability Engineering (SRE) Lead - 2770

Role overview

Qualifications

  • 8-10+ years of experience in SRE/Observability/DevOps with leadership responsibilities.
  • Hands-on experience with OpenTelemetry for distributed tracing and observability instrumentation; strong exposure to APM tools (New Relic, Datadog, AppDynamics, Dynatrace).
  • Proficiency with Terraform and other IaC practices, plus cloud platforms (AWS, GCP, Azure).
  • Experience with automation/configuration management (Ansible, Chef, Puppet), CI/CD pipelines, and Kubernetes (Docker, Helm).

Responsibilities

  • Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business goals and technical requirements.
  • Design and implement monitoring and observability solutions, collaborating with engineering teams to define standards and best practices.
  • Drive automation strategies and IaC initiatives (Terraform) while coordinating with cloud/infrastructure teams to ensure scalable, secure deployments and robust monitoring/logging pipelines.
  • Mentor junior engineers and analysts, and collaborate with product management, sales, and pre-sales during solution design and customer engagements, including vendor and partner engagements when applicable.

About the company

Xideral logo

Xideral

IT Services & IT Consulting

We are a proudly Mexican corporation, with 20 years of experience and operations in Mexico, LATAM and Canada. We have a highly qualified team that allows us to create tailored solutions, implementations and provide services in the IT area, with the aim of achieving the goals of our customers and support their growth. We are committed to the educational and economic development of our country, so we design employment de-centralization programs that help reduce the migration of young people from disadvantaged regions, which creates a chain of welfare for the communities.

Company details

IndustryIT Services & IT Consulting
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Seeking Experienced Site Reliability Engineering (SRE) / Lead Engineer for Exciting Projects Remote in Guadalajara, Jalisco

 

We are looking for skilled Site Reliability Engineering (SRE) / Lead Engineer with a minimum of 8 years of experience to join a dynamic team within a leading organization. This role must have deep expertise in Application Performance Monitoring (APM), Infrastructure as Code (IaC), automation, and distributed tracing using OpenTelemetry.

As a SRE lead, he will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. 

 

Key Responsibilities:

·     -Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business goals and technical requirements.

·     -Design and implementation of monitoring and observability solutions, collaborating with engineering teams to define standards and best  practices.

·      -Manage Infrastructure as Code (IaC) initiatives using Terraform, coordinating with cloud and infrastructure teams to ensure scalable and secure deployments.

·      -Drive automation strategies for monitoring, alerting, and logging pipelines, focusing on process improvements and operational efficiency.

·       -Develop and maintain comprehensive observability roadmaps, including distributed tracing, logging, and metrics collection strategies.

·      -Collaborate with product management, sales, and pre-sales teams to provide technical expertise and support during solution design and customer engagements.

·    -Lead cross-functional teams to enhance CI/CD pipelines and deployment reliability, ensuring smooth integration of observability tools and practices.

·      -Engage with vendors and strategic partners to evaluate, select, and integrate observability and monitoring solutions, ensuring alignment with organizational needs and fostering strong collaborative relationships.

·      -Mentor and develop junior engineers and analysts, fostering a culture of reliability, observability, and operational excellence.

 

Technical Skills Required:

·       -  8-10+ years of experience in SRE, Observability, or DevOps roles, with leadership responsibilities.

·        - Hands-on experience with OpenTelemetry for distributed tracing and observability instrumentation.

·         -Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace.

·         -Strong proficiency in Infrastructure as Code (IaC) using Terraform.

·         -Solid understanding of cloud platforms including AWS, GCP, or Azure.

·         -Experience with automation/configuration management tools like Ansible, Chef, or Puppet.

·         -Deep knowledge of CI/CD pipelines and tools such as GitHub Actions, Jenkins, or Azure DevOps.

·         -Experience managing Kubernetes and containerized environments (Docker, Helm).

·         -Familiarity with log aggregation and analysis platforms like ELK Stack or Splunk.

·         -Excellent leadership, communication, and collaboration skills.

 

Location & Schedule:

  • Remote work in Guadalaraja, Jaliso
  • Work hours Monday to Friday, 09:00 – 18:00
  • Advanced English skills are mandatory

 

Benefits:

·         Attractive Salary + Premium Benefits

·         Performance bonuses,  grocery coupons, and savings are found.

·         Aguinaldo,  premium vacations,  and vacations paid

·         SGMM Medical insurance, family, and  Life insurance.

 

Candidates must include their compensation expectations in their applications and resumes in English.

 

Interested? Apply now through this link

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Xideral

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.