Logo for Compass.uol

Site Reliability Engineer (SRE) | AWS | Kubernetes | Databricks | Sênior (Remote)

Role overview

Qualifications

  • Solid experience as Site Reliability Engineer (SRE), DevOps Engineer or Platform Engineer
  • Experience in managing Kubernetes environments in production
  • Experience with cloud infrastructure, preferably AWS
  • Knowledge of infrastructure automation using Terraform and Infrastructure as Code (IaC)

Responsibilities

  • Ensure the reliability, availability, and performance of production systems and applications
  • Design, implement, and evolve automations for deployment, monitoring, scalability, and operation of the platform
  • Administer and evolve Kubernetes environments, ensuring stability and high availability
  • Respond to critical incidents, conducting root cause analysis (RCA) and implementing preventive actions

About the company

Compass.uol logo

Compass.uol

Digital Transformation Consulting

Compasso UOL is a Brazilian technology company owned by the UOL Group, which offers technology services, and provides state-of-the-art solutions, contributing to the digital transformation of its customers so they can become leaders in its sectors of activity. Compasso UOL is a company that cultivates people's talent and use state-of-the-art technologies such as Agile Development, Multicloud, Data&Analytics, Cyber Security, Artificial Intelligence/Machine Learning, APIs/Microservices, IoT and others, providing technology and knowledge that help customers in building digital solutions that enable the transformation and evolution of their business.

Company details

Company typeXLarge
IndustryDigital Transformation Consulting
Company size5001 - 10000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

JOB DESCRIPTION


.


RESPONSIBILITIES AND ASSIGNMENTS


  • Garantir a confiabilidade, disponibilidade e desempenho de sistemas e aplicações em produção;
  • Projetar, implementar e evoluir automações para deploy, monitoramento, escalabilidade e operação da plataforma;
  • Administrar e evoluir ambientes Kubernetes, assegurando estabilidade e alta disponibilidade;
  • Responder a incidentes críticos, conduzindo análises de causa raiz (RCA) e implementando ações preventivas;
  • Definir, acompanhar e evoluir indicadores de confiabilidade, como SLIs, SLOs e SLAs;
  • Gerenciar e otimizar pipelines de CI/CD, promovendo entregas contínuas e seguras;
  • Implementar e manter infraestrutura como código (IaC), garantindo padronização e automação dos ambientes;
  • Desenvolver e fortalecer práticas de observabilidade, monitoramento, métricas e alertas;
  • Colaborar com equipes de desenvolvimento na construção de aplicações resilientes, escaláveis e de alta performance;
  • Documentar procedimentos operacionais, planos de contingência e boas práticas de operação;
  • Atuar na melhoria contínua da plataforma, reduzindo falhas recorrentes e aumentando a confiabilidade dos serviços.

REQUIREMENTS AND QUALIFICATIONS


  • Experiência sólida como Site Reliability Engineer (SRE), DevOps Engineer ou Platform Engineer;
  • Vivência em administração de ambientes Kubernetes em produção;
  • Experiência com infraestrutura em cloud, preferencialmente AWS;
  • Conhecimento em automação de infraestrutura utilizando Terraform e Infrastructure as Code (IaC);
  • Experiência com pipelines CI/CD, processos de integração e entrega contínua;
  • Vivência com ferramentas de observabilidade, monitoramento e gestão de incidentes;
  • Experiência em troubleshooting de aplicações distribuídas e ambientes críticos de produção;
  • Experiência com Databricks e Apache Spark;
  • Experiência com AWS CodePipeline;
  • Conhecimento em bancos de dados SQL, Data Warehouse e modelagem de dados;
  • Experiência com Amazon EC2, Amazon S3 e AWS Lambda;
  • Conhecimento em arquitetura cloud e ambientes Microsoft Azure;
  • Conhecimento em práticas de alta disponibilidade, escalabilidade, resiliência e recuperação de incidentes;
  • Capacidade de atuar de forma analítica na identificação de causas raiz e implementação de melhorias contínuas;
  • Diferencial: Certificações Databricks.



Become a Compasser, be part of AI/R.


Compass UOL is a global firm and part of the AI Revolution Company, together transforming organizations using Artificial Intelligence, Generative AI, and other of today’s most advanced technologies.


We equip our team with proprietary and external AI-driven tools to design and build digital-native platforms, integrating cutting-edge technologies and enabling companies to innovate, transform their businesses, and drive success in their markets.

To achieve this, we attract and develop the best talent, creating opportunities that enhance people’s lives and highlight the positive impact of disruptive technologies.

We empower borderless talent and promote knowledge and opportunities in the latest market trends, driving significant personal and professional growth.

Join us and be part of the AI-driven revolution.


Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Compass.uol

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.