Logo for Reka AI

Member of Technical Staff- Data Intelligence

Role overview

Qualifications

  • Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems
  • Comfortable moving between research questions and production engineering; ability to dig into data, run analyses, and ship reliable systems
  • Demonstrated research experience with data compositions, quality, and dataset releases
  • Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)

Responsibilities

  • Define what good data means for models, including quality metrics, validation checks, and acceptance thresholds
  • Explore open source datasets and create internal ones most suitable to build fundamental World Models
  • Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data
  • Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs

About the company

Reka AI logo

Reka AI

Artificial Intelligence & Machine Learning Services

AI Research and Product company with a mission to build amazing generative models and advance AI research.

Company details

Company typeStartup
IndustryArtificial Intelligence & Machine Learning Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.

What you’ll do

  • Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds

  • Explore open source datasets and create internal ones most suitable to build fundamental World Models

  • Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.

  • Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs

  • Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction

  • Track and optimize throughput, storage, and compute utilization across pipelines and related assets

What we’re looking for

  • Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems

  • Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems

  • Demonstrated research experience with data compositions, quality, and dataset releases

  • Ability to design and execute experiments with convincing unbiased outcomes

  • Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)

  • Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)

  • Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale

  • Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers

  • Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Reka AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.