Logo for Avra

Member of Technical Staff | Inference Platform

Role overview

Qualifications

  • Experience running model serving or large-scale batch compute on Kubernetes
  • Experience building Kubernetes controllers or operators
  • Skill at profiling and optimizing data-heavy Python pipelines
  • A clear sense of cost: you treat compute efficiency as a product feature

Responsibilities

  • Evolve Sophos, our online and batch inference runtime, built on Kubernetes
  • Run large batch inference on ephemeral jobs, with multi-dimensional admission control (CPU, memory, GPU) through Kueue
  • Build and extend the Sophos controller and its Kubernetes custom resources
  • Optimize each model's inference engine and feature processing, using vectorized, columnar operations

Key facts

  • Remote from: Brazil
  • Full time
  • Senior (5-10 years)
  • English

Hard skills

About the company

Avra logo

Avra

Artificial Intelligence & Machine Learning Services

A Avra desenvolve modelos fundacionais baseados em large knowledge graphs para entender como o mercado de pequenas e médias empresas funciona.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.

In this role, you'll join the Platform team to own where our models execute. Customers consume our models through large batches of millions of records and through real-time APIs, and they make business decisions on every response. You'll run governed model releases reliably and efficiently — in our cloud and on customer-hosted Kubernetes — and make inference fast, predictable, and cheap enough to serve both enterprise and mid-market customers.

What you'll do

  • Evolve Sophos, our online and batch inference runtime, built on Kubernetes.

  • Run large batch inference on ephemeral jobs, with multi-dimensional admission control (CPU, memory, GPU) through Kueue.

  • Build and extend the Sophos controller and its Kubernetes custom resources.

  • Optimize each model's inference engine and feature processing, using vectorized, columnar operations.

  • Serve graphs and data efficiently from Lance-based storage.

  • Own execution of training, post-training, and fine-tuning jobs, in our cloud and in customer dataplanes.

  • Drive autoscaling, GPU serving, performance, and cost optimization, with telemetry for every model we run.

  • Solve open problems such as deterministic job sizing, checkpointing and recovery for batch runs, per-customer encryption and isolation, resilience to difficult input files, and automatic profiling when a new model is accepted.

How we measure success

  • 99.9% serving availability.

  • p95/p99 latency for online inference and throughput for batch.

  • Cost per prediction and per training job.

  • GPU utilization: paid capacity versus capacity actually used.

  • Training and batch jobs that finish on time and succeed without manual retries.

What we're looking for

  • Experience running model serving or large-scale batch compute on Kubernetes.

  • Experience building Kubernetes controllers or operators.

  • Skill at profiling and optimizing data-heavy Python pipelines.

  • A clear sense of cost: you treat compute efficiency as a product feature.

  • Production-quality code and reviews, and a willingness to operate what you build.

Nice to have

  • Ray, Ray Serve, or KubeRay in production.

  • Kueue or other batch scheduling and admission-control systems.

  • GPU serving and performance optimization.

  • Arrow, Parquet, Lance, or other columnar formats.

  • Shipping software to customer-hosted Kubernetes.

  • GCP/AWS and GKE/EKS, and financial services or regulated environments.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Avra

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.