Logo for Avra

Member of Technical Staff | Observability & Reliability

Role overview

Qualifications

  • Deep experience with OpenTelemetry and observability backends
  • Hands-on practice with SLOs, error budgets, actionable alerting, and incident management
  • Strong experience with Kubernetes and infrastructure as code (Terraform / Helm)
  • Experience operating software in environments you don't fully control

Responsibilities

  • Evolve our observability stack for logs, metrics, traces, and alerting
  • Ensure all dataplanes report their active release, health, heartbeat, logs, metrics, and usage
  • Bring telemetry into customer clusters with outbound-only connections
  • Monitor health of deployment and runtime agents

Key facts

  • Remote from: Brazil
  • Full time
  • Senior (5-10 years)
  • Technical Support Manager
  • English

Hard skills

About the company

Avra logo

Avra

Artificial Intelligence & Machine Learning Services

A Avra desenvolve modelos fundacionais baseados em large knowledge graphs para entender como o mercado de pequenas e médias empresas funciona.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.

In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.

What you'll do

  • Evolve our observability stack for logs, metrics, traces, and alerting.

  • Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.

  • Bring telemetry into customer clusters within a model where agents only make outbound connections.

  • Detect drift between the desired state and what's actually running in each environment.

  • Monitor the health of our deployment and runtime agents.

  • Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.

  • Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.

  • Reduce telemetry cost: less redundant data, more useful signal.

How we measure success

  • 99.9% serving availability, with incidents trending down.

  • MTTR, including on-premise incidents.

  • Near-zero drift between desired and actual state.

  • All agents active and reporting, across every dataplane.

What we're looking for

  • Deep experience with OpenTelemetry and observability backends.

  • Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.

  • Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).

  • Experience operating software in environments you don't fully control.

  • Production-quality code and reviews, and a willingness to operate what you build.

Nice to have

  • Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).

  • GCP or GKE, AWS or EKS.

  • ML multi-node/multi-cluster workloads in production.

  • Financial services or regulated environments.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Technical Support Manager Related jobs

Other jobs at Avra

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.