Logo for Software Mind

[8SN] Senior Site Reliability / Production Support Engineer

Role overview

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations
  • Strong hands-on experience with Kubernetes in production environments
  • Experience supporting cloud-native applications
  • Strong knowledge of Splunk for log analysis and debugging

Responsibilities

  • Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes
  • Monitor application health, availability, and performance
  • Investigate and resolve production incidents using logs, monitoring, and debugging tools
  • Collaborate with software engineers to improve service reliability and operational efficiency

About the company

Software Mind logo

Software Mind

Software Mind is a global digital transformation partner with operations throughout Europe, the US and LATAM. Driven by tech and empowered by people, we provide companies with software engineers and autonomous, cross-functional development teams who manage software life cycles from ideation to release and beyond. For over 20 years we’ve been enriching organizations with the talent they need to boost scalability, drive dynamic growth and bring disruptive ideas to life. Our top-notch engineering teams combine ownership with leading technologies, including cloud, AI, data science and embedded software to accelerate digital transformations and boost software delivery. A culture, driven by trust, that embraces openness, craves more and acts with respect enables our experts to create evolutive solutions that support scale-ups, unicorns and enterprise-level companies around the world.

Company details

Company typeLarge
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
 

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

#LI-DNI

Job Description

About the Role

We are looking for a Senior Site Reliability / Production Support Engineer to support the deployment, operations, and ongoing reliability of a production UI service running on Kubernetes.

This role is focused on maintaining highly available cloud-native applications, troubleshooting production issues, and improving operational excellence. You will work closely with engineering teams to monitor service health, investigate incidents, and ensure reliable service delivery.

While this role supports a UI-based service, it is not a frontend development position. Basic knowledge of Web Components is sufficient to perform first-level debugging when necessary.
 

What You'll Do

  • Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes.
  • Monitor application health, availability, and performance.
  • Investigate and resolve production incidents using logs, monitoring, and debugging tools.
  • Perform log analysis using Splunk to identify root causes and troubleshoot service issues.
  • Collaborate with software engineers to improve service reliability and operational efficiency.
  • Participate in incident response and production support activities.
  • Assist with first-level debugging of UI-related issues involving Web Components.
  • Contribute to continuous improvements in automation, monitoring, and operational processes.
  • Support CI/CD pipelines and cloud-native deployment practices.

Qualifications

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.
  • Strong hands-on experience with Kubernetes in production environments.
  • Experience supporting cloud-native applications.
  • Experience monitoring production systems and troubleshooting complex incidents.
  • Strong knowledge of Splunk for log analysis and debugging.
  • Experience working in Linux environments.
  • Understanding of networking fundamentals and distributed systems.
  • Experience collaborating with software engineering teams to resolve production issues.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent written and spoken English (B2+).

Additional Information

Preferred Qualifications

  • Experience with CI/CD pipelines.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Familiarity with container technologies such as Docker.
  • Exposure to observability tools (Prometheus, Grafana, OpenTelemetry, etc.).
  • Basic understanding of Web Components and frontend architecture.
  • Experience supporting high-availability enterprise SaaS platforms.
  • Knowledge of infrastructure automation or Infrastructure as Code (Terraform, Helm, Ansible, etc.) is a plus.

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Software Mind

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.