Logo for Netlution GmbH

Senior Platform Operations Engineer / Site Reliability Engineer (m/w/d)

Role overview

Qualifications

  • Multiple years of experience in Platform Operations, Site Reliability Engineering, System Engineering or IT Operations
  • Very good knowledge in Linux-based environments
  • Understanding of Kubernetes and containerized platforms
  • Experience in monitoring, alerting, and observability environments

Responsibilities

  • Operation, maintenance, and development of monitoring and observability platforms
  • Administration and optimization of Prometheus, Grafana, and OpenSearch / ELK
  • Monitoring productive system landscapes and continuous improvement of monitoring and alerting concepts
  • Handling incidents and supporting the recovery of critical services

Key facts

  • Remote from: Germany
  • Full time
  • Senior (5-10 years)
  • Site Reliability Engineer (SRE)
  • German

Hard skills

Other skills

  • Analytical Thinking
  • Problem Solving

About the company

Netlution GmbH logo

Netlution GmbH

IT Services & IT Consulting

The n_Story In 2001 the story of Netlution began. Over the years the company has established itself as an IT infrastructure specialist with adaptive IT services and consulting services in the areas of data center, end user and cloud computing of complex IT environments in the area. As a partner of corporations in the enterprise environment, we look behind the scenes of the IT departments of large corporations, which are active in a wide variety of areas. 180 employees shape our motto: our people make the difference, because they make our services unique and that makes us proud. Our success and growth is shaped by each team member, which is why we support each individual in his or her personal and professional development. 16 values have been developed by our team members in a participatory way. We live these values every day in order to realize our visions. For us, teamwork means devoting ourselves to exciting projects together and mastering challenges with team spirit. Our sense of community is also shaped by events where we meet regularly to spend time together in a relaxed atmosphere. Find out more about us at www.our-people-make-the-difference.de!

Company details

IndustryIT Services & IT Consulting
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Über unseren Kunden


Aufgabenspektrum

Sichere und betreibe moderne Plattformen auf Enterprise-Niveau.

Du begeisterst dich für hochverfügbare Plattformen, Monitoring-Lösungen und den stabilen Betrieb komplexer Kubernetes-Umgebungen? Dann suchen wir genau dich.

Als Senior Platform Operations Engineer (m/w/d) übernimmst du eine zentrale Rolle im Betrieb und der Weiterentwicklung moderner Plattform- und Observability-Lösungen. Dein Fokus liegt nicht auf klassischer Softwareentwicklung, sondern auf der Sicherstellung von Stabilität, Verfügbarkeit und Performance geschäftskritischer Systeme.

Du arbeitest an den Themen Monitoring, Observability, Incident Management, Kubernetes Operations sowie Automation im Betriebsumfeld und sorgst gemeinsam mit unserem Team für einen zuverlässigen Plattformbetrieb.

Deine Aufgaben

  • Betrieb, Betreuung und Weiterentwicklung von Monitoring- und Observability-Plattformen
  • Administration und Optimierung von Prometheus, Grafana sowie OpenSearch / ELK
  • Überwachung produktiver Systemlandschaften und kontinuierliche Verbesserung von Monitoring- und Alerting-Konzepten
  • Bearbeitung von Incidents sowie Unterstützung bei der Wiederherstellung kritischer Services
  • Durchführung von Root Cause Analysen und nachhaltige Beseitigung von Störungsursachen
  • Unterstützung bei Major Incidents und Koordination technischer Lösungsmaßnahmen
  • Betrieb und Optimierung containerisierter Plattformen auf Basis von Kubernetes
  • Unterstützung von CI/CD-Prozessen mit Jenkins und ArgoCD
  • Erstellung, Pflege und Weiterentwicklung von Runbooks, Betriebsprozessen und technischer Dokumentation
  • Automatisierung wiederkehrender Betriebsaufgaben
  • Teilnahme an Rufbereitschaften und Schichtmodellen in einer 24x7-Betriebsorganisation


Erfahrungen

Das bringst du mit:

Must-have

  • Mehrjährige Erfahrung im Bereich Platform Operations, Site Reliability Engineering, System Engineering oder IT Operations
  • Bereitschaft zur Durchführung bzw. Vorliegen einer Sicherheitsüberprüfung SÜ2
  • Sehr gute Kenntnisse in Linux-basierten Umgebungen
  • Verständnis für Kubernetes und containerisierte Plattformen
  • Praxiserfahrung mit:
    • Prometheus
    • Grafana
    • ELK Stack oder OpenSearch
    • Elasticsearch oder OpenSearch
  • Erfahrung im Monitoring, Alerting und Observability-Umfeld
  • Gute Kenntnisse von Netzwerkgrundlagen und Kommunikationsprotokollen
  • Erfahrung im Umgang mit REST APIs
  • Sicherer Umgang mit Git
  • Analytische Herangehensweise bei Fehleranalysen und Störungsbehebung
  • Gute Englischkenntnisse in Wort und Schrift
  • Bereitschaft zur Teilnahme an Rufbereitschaften und Schichtbetrieb
Nice-to-have
  • Erfahrung mit ArgoCD
  • Kenntnisse in Jenkins
  • Erfahrung mit Helm
  • Bash-Scripting
  • Python für Betriebsautomation und Operational Excellence
  • Erfahrung im Bereich Site Reliability Engineering (SRE)
  • Kenntnisse moderner Cloud- oder Plattformarchitekturen

Benefits

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Netlution GmbH

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.