Logo for Anyone AI

Machine Learning Engineer – ML Evaluation & Experiment Design

Role overview

Qualifications

  • 3+ years of hands-on applied machine learning experience
  • Strong experience with ML experiment design and model selection
  • Ability to identify data leakage, label noise, distribution shift, and feature leakage
  • Strong understanding of ML evaluation metrics and debugging ML workloads

Responsibilities

  • Reviewing ML challenges and determining whether they are well designed and technically solvable
  • Evaluating whether datasets contain meaningful and learnable signals
  • Identifying unintended shortcuts or artifacts in synthetic datasets
  • Verifying reproducibility across the complete data → model → evaluation pipeline

Key facts

Hard skills

Other skills

  • Analytical Skills

About the company

Anyone AI logo

Anyone AI

E-Learning / EdTech

We invest in talent from Latam to bridge the talent gap in AI. Join our AI community: www.anyoneai.com We are AI / ML experts and second-time entrepreneurs, members of the founding team at Deep Vision AI (acquired in early 2020). We've worked with many Fortune 500 companies completing multiple projects in the early days of AI while leveraging remote talent from LatAm. We are VC-backed from day 1 by top global investors like GFC -Global Founders Capital- (investors in Facebook, LinkedIn, Slack, Canva, Trivago, etc), Canvas Ventures (early investor in Coursera), Latitud Fund (the largest community of angel investors for LatAm including investments in QuintoAndar, La Haus, Clara, Platzi, OnTop, Pomelo, etc), among other investors.

Company details

IndustryE-Learning / EdTech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Experiment design and model selection

  • Small and synthetic datasets

  • Data quality and preprocessing

  • Distribution shift and data contamination

  • Label noise and feature leakage

  • Model evaluation and metric selection

  • Hyperparameter tuning

  • Train / validation / test methodology

  • Reproducibility and deterministic pipelines

  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

  • 3+ years of hands-on applied machine learning experience

  • Strong experience with:

    • ML experiment design

    • Model selection

    • Hyperparameter tuning

    • Model evaluation

    • Data preprocessing and validation

  • Strong understanding of train, validation, and test splits

  • Ability to identify:

    • Data leakage

    • Label noise

    • Distribution shift

    • Spurious correlations

    • Feature leakage

    • Data contamination

  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations

  • Strong understanding of ML evaluation metrics and when different metrics are appropriate

  • Experience debugging ML workloads across CPU and GPU environments

  • Ability to analyze technical problems and provide clear written feedback

Nice to Have

  • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions

  • Experience designing benchmark datasets or ML challenges

  • Background in data-centric AI or dataset quality

  • Experience with synthetic data generation and validation

  • Familiarity with statistical testing, confidence intervals, and effect sizes

  • Experience with ML evaluation pipelines, RLHF, or AI model evaluation

  • Experience developing ML curricula or technical assessments

  • Understanding of common ML failure modes such as:

    • Shortcut learning

    • Spurious correlations

    • Goodhart’s Law

    • Simpson’s paradox

    • Metric gaming

What You’ll Be Responsible For

  • Reviewing ML challenges and determining whether they are well designed and technically solvable

  • Evaluating whether datasets contain meaningful and learnable signals

  • Identifying unintended shortcuts or artifacts in synthetic datasets

  • Determining whether tasks require genuine diagnosis of the underlying ML problem

  • Reviewing evaluation metrics and improvement thresholds

  • Detecting metric gaming, data leakage, and evaluation flaws

  • Verifying reproducibility across the complete data → model → evaluation pipeline

  • Assessing whether challenge difficulty is appropriately calibrated

  • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Machine Learning Engineer Related jobs

Other jobs at Anyone AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.