Logo for Avra

Member of Technical Staff | ML Systems

Role overview

Qualifications

  • Strong systems engineering skills and production-quality Python
  • Experience with distributed training (e.g., Ray, PyTorch distributed)
  • Experience with columnar data formats and large-scale data materialization
  • Familiarity with ML lifecycle tooling: experiment tracking, model registries, evaluation, and reproducibility

Responsibilities

  • Build CUDA kernels and compute primitives for training and serving graph neural networks (GNNs)
  • Evolve Monad, our sampler and distributed-training library, including neighbor sampling and training performance
  • Specify our binary data formats (Lance, Arrow, CSR/CSC), and own materializations and feature backfills for training and evaluation
  • Define and run release gates, so every model running in production, batch, or on-premise maps to a governed release

Key facts

  • Remote from: Brazil
  • Full time
  • Senior (5-10 years)
  • Technical Support Manager
  • English

Hard skills

Other skills

  • Governance

About the company

Avra logo

Avra

Artificial Intelligence & Machine Learning Services

A Avra desenvolve modelos fundacionais baseados em large knowledge graphs para entender como o mercado de pequenas e médias empresas funciona.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.

In this role, you'll join our ML Systems team, which owns Avra's ML core and the governance of every model we ship. Research produces candidate models and evidence; you build the reliable path from data and training to a governed, reproducible release that can run in our cloud or in any customer environment. ML Systems is an internal platform: its users are our researchers and platform engineers, and its success is measured by the leverage it creates for them.

What you'll do

  • Build CUDA kernels and compute primitives for training and serving graph neural networks (GNNs).

  • Evolve Monad, our sampler and distributed-training library, including neighbor sampling and training performance.

  • Specify our binary data formats (Lance, Arrow, CSR/CSC), and own materializations and feature backfills for training and evaluation.

  • Define data contracts and consumption requirements with the teams that build our customer and proprietary datasets.

  • Build and operate experiment tracking, checkpoints, and evaluation infrastructure, with reproducibility by default.

  • Own the model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models.

  • Define and run release gates, so every model running in production, batch, or on-premise maps to a governed release.

  • Make it possible to audit exactly which data, code, configuration, and evidence produced each release.

How we measure success

  • Time-to-experiment: how quickly a researcher goes from a hypothesis to materialized data, compute, and tracking.

  • Time-to-governed-release: how quickly a validated candidate becomes an authorized release.

  • Training throughput per GPU on our foundation model training runs.

  • 100% of production models with complete release records and lineage — no ad hoc models in any environment.

  • Every release reproducible from its registered data, code, and configuration.

What we're looking for

  • Strong systems engineering skills and production-quality Python.

  • Experience with distributed training (e.g., Ray, PyTorch distributed) and multi-node GPU workloads.

  • Experience with columnar data formats and large-scale data materialization.

  • Familiarity with ML lifecycle tooling: experiment tracking, model registries, evaluation, and reproducibility.

  • A product mindset: you treat an internal platform as a product with real users. You don't need to be a data scientist.

Nice to have

  • CUDA kernel development or GPU performance optimization.

  • Graph neural networks or graph sampling at scale.

  • Lance, Arrow, or other columnar/indexed storage formats.

  • Multi-cloud GPU compute (e.g., SkyPilot).

  • Model governance or audit requirements in financial services or other regulated environments.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Technical Support Manager Related jobs

Other jobs at Avra

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.