Logo for Fundamental

Evaluations Team Lead

Role overview

Qualifications

  • Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.
  • Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.
  • Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.
  • Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.

Responsibilities

  • Build a shared evaluation platform for engineering, research, and Applied AI.
  • Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.
  • Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.
  • Hire and develop the team, set priorities, and stay hands-on with code and experimental design.

Key facts

Hard skills

Other skills

  • People Management

About the company

Fundamental logo

Fundamental

Artificial Intelligence & Machine Learning Services

For decades companies have relied on archaic tools to inform decisions and make bets on the future. Until now. Fundamental empowers businesses to turn gambles into guarantees and determine their future with far greater accuracy than ever before. Built by DeepMind alumni and trusted by Fortune 100 enterprises, NEXUS is our most powerful Large Tabular Model (LTM). By revealing the hidden language of tables, NEXUS unlocks trillions of dollars of value by giving businesses the Power to Predict™.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Fundamental

Fundamental is an AI research lab pioneering the future of enterprise decision-making. Our flagship model, NEXUS is the world's most powerful Large Tabular Model (LTM) - purpose-built for the structured records that contain trillions of dollars in business value. With $275m in funding from leading investors and trusted by Fortune 100 companies, Fundamental is giving businesses the Power to Predict.

At Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.

About the role

NEXUS is already in production, and we are working to substantially improve its predictive quality, latency, and cost efficiency. You will lead the team that owns how NEXUS is measured, building the shared evaluation platform that engineering, research, and Applied AI all rely on: consistent benchmarks, curated datasets, and standards for metrics, data splits, leakage prevention, and benchmark contamination that hold up under scrutiny.

Research needs to trust results before committing compute to it, engineering needs regressions caught before they reach production, and when a customer's own data science team benchmarks NEXUS against their own models, your platform is what Applied AI relies upon. You will benchmark NEXUS against competing approaches, using fair tuning budgets, data access, and latency measurement protocols, and turn what you find into research priorities and release recommendations.

This is a player-coach role: you will hire and manage a small team while staying hands-on with the code, the experiment design, and the methodology yourself. There is no evaluation function to inherit here - what you build becomes the standard the rest of the company measures NEXUS against.

Key responsibilities

  • Build a shared evaluation platform for engineering, research, and Applied AI. Continuously add and maintain models and curated datasets so teams can run benchmarks and investigate results independently.

  • Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.

  • Build reproducible pipelines with versioned inputs and artifacts, and integrate regression checks into research and release workflows.

  • Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.

  • Support Applied AI’s customer POC evaluations with tooling, methodological guidance, and analysis.

  • Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics.

  • Turn findings into research priorities, release recommendations, and evidence-backed customer improvement plans.

  • Hire and develop the team, set priorities, and stay hands-on with code and experimental design.

Must have

  • Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.

  • Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.

  • Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits.

  • Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.

  • Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement.

  • Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.

  • Clear written and spoken communication with researchers, engineers, and customer data scientists, including the willingness to challenge claims the evidence does not support.

Nice to have

  • Experience evaluating tabular foundation models or AutoML systems.

  • Experience measuring how model optimisations affect predictive quality and inference performance.

  • Experience with relational or multi-table data, and Snowflake or Databricks environments.

  • Experience evaluating automated or agent-driven ML workflows, including failures that aggregate metrics can hide.

Benefits

  • Competitive compensation with salary and equity

  • Comprehensive health coverage for you and your dependents

  • Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys

  • Relocation support for employees moving to join the team in one of our office locations

  • A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Fundamental

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.