Logo for 99x

Observability Engineer - Site Reliability Engineer

Role overview

Qualifications

  • 3+ years of experience as an SRE, Observability Engineer, or equivalent role
  • Practical experience with OpenTelemetry, or similar instrumentation tools
  • Experience in Kubernetes, Helm, Terraform, and ArgoCD
  • Fluency in English

Responsibilities

  • Design, implement, and maintain observability solutions covering metrics, logs, traces, and RUM
  • Work with tools such as Grafana Cloud, Tempo, Loki, Mimir, Alloy, and OpenTelemetry
  • Build reliable alerting and monitoring pipelines based on SLOs/SLAs, focusing on low-maintenance automation
  • Collaborate with development and operations teams to embed observability by design into the software lifecycle

Key facts

Hard skills

Other skills

  • Collaboration
  • Problem Solving

About the company

99x logo

99x

Software Development

Headquartered in Norway, 99x is a technology company co-creating well-engineered, innovative digital products for the Scandinavian market. Its expertise has been proven through a portfolio of over 150 impactful global digital products developed since 2004, together with leading Independent Software Vendors (ISVs). 99x employs over 500 technology and product specialists, who are high achievers, creative thinkers and team players. The company is one of Asia’s Best Workplaces for 2022 and has been named a Best Workplace in Sri Lanka for 10 consecutive years.

Company details

Company typeSME
IndustrySoftware Development
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

This is a remote position.

We are looking for a Site Reliability Engineer (SRE) with strong expertise in Observability Engineering to join our team. This role is pivotal to ensuring the reliability, visibility, and performance of our platforms and services. The ideal candidate will have hands-on experience with the Grafana Stack (Tempo, Loki, Mimir, Alloy), knowledge in Java development, a strong SRE mindset, and a passion for automation, scalability, and ownership.

You’ll be joining a motivated, cross-functional team responsible for implementing and scaling our new observability stack that is being built to be a platform for observability for the whole company. Your contributions will directly impact system performance, user experience, and operational cost efficiency.



Requirements

Responsibilities

·       Design, implement, and maintain observability solutions covering metrics, logs, traces, and RUM.

·       Work with tools such as Grafana Cloud, Tempo, Loki, Mimir, Alloy, and OpenTelemetry.

·       Build reliable alerting and monitoring pipelines based on SLOs/SLAs, focusing on low-maintenance automation.

·       Ensure the health and integrity of observability data flows from instrumentation to dashboards.

·       Collaborate with development and operations teams to embed observability by design into the software lifecycle.

·       Define and promote best practices and standards for observability across the organization.

·       Support the modernization of observability by replacing and evolving legacy monitoring and alerting solutions.

·       Monitor observability-related costs and contribute to FinOps efforts by identifying optimization opportunities.

Requirements

Must-have:

·       3+ years of experience as an SRE, Observability Engineer, or equivalent role.

·       Practical experience with OpenTelemetry, or similar instrumentation tools.

·       Experience in Kubernetes, Helm, Terraform, and ArgoCD.

·       Experience designing and managing telemetry pipelines (metrics/logs/traces), exporters, and sidecars.

·       Product-oriented mindset with a bias for automation and a “you build it, you run it” culture

·       Fluency in English.

Nice-to-have:

·       Knowledge of APM and distributed tracing solutions.

·       Experience with FinOps practices applied to observability.

·       Hands-on involvement in replacing legacy monitoring stacks.

·       Experience with Cloud environments (Azure preferred)

·       Contributions to open-source observability tools.

·       Knowledge in Java development and applications instrumentation

·       Expertise in performance monitoring, alerting, dashboarding, and root cause analysis.



Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at 99x

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.