Logo for Anyone AI

Research Scientist (Remote/US/LATAM)

Role overview

Qualifications

  • Research background in ML evaluation or benchmarking
  • Deep LLM benchmarking expertise
  • Fluency with how frontier models are measured
  • Proven ability to hold a team or expert pool to a rigorous standard

Responsibilities

  • Evaluation research to create original evaluation designs
  • Build evaluation packages with subject-matter experts
  • Recruit, calibrate, and review expert pools across various domains
  • Act as a technical point of contact for labs and understand their measurement needs

About the company

Anyone AI logo

Anyone AI

E-Learning / EdTech

We invest in talent from Latam to bridge the talent gap in AI. Join our AI community: www.anyoneai.com We are AI / ML experts and second-time entrepreneurs, members of the founding team at Deep Vision AI (acquired in early 2020). We've worked with many Fortune 500 companies completing multiple projects in the early days of AI while leveraging remote talent from LatAm. We are VC-backed from day 1 by top global investors like GFC -Global Founders Capital- (investors in Facebook, LinkedIn, Slack, Canva, Trivago, etc), Canvas Ventures (early investor in Coursera), Latitud Fund (the largest community of angel investors for LatAm including investments in QuintoAndar, La Haus, Clara, Platzi, OnTop, Pomelo, etc), among other investors.

Company details

IndustryE-Learning / EdTech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Research Scientist, Evaluations

Anyone AI Labs โ€” Human Data Division
Reports to: CEO ยท Remote / LatAm / US

The role

You will own how Anyone AI measures frontier model capability. This is a research role at heart: you decide what a good evaluation is, design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal โ€” grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review.

Responsibilities

  • Evaluation research. Turn public benchmarks and eval targets into original evaluation designs. Own the hard questions: construct validity, discrimination, headroom, and contamination.

  • Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).

  • Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.

  • Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.

  • Delivery. Turn lab requests into winning sample packages, then own pilots end to end. Nothing ships before it's lab-ready.

What we're looking for

  • Research background in ML evaluation or benchmarking โ€” published/open benchmarks, eval research, or equivalent hands-on work labs relied on.

  • Deep LLM benchmarking expertise, with real strength in code-model evaluation.

  • Fluency with how frontier models are measured: rubrics, pass rates, headroom, contamination, and what makes a task discriminate a model.

  • Proven ability to hold a team or expert pool to a rigorous standard.

  • Fluent English. Spanish a nice to have.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
ยท

Researcher Related jobs

Other jobs at Anyone AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.