Logo for Anyone AI

Human Data Evals Lead (Remote/US/LATAM)

Role overview

Qualifications

  • 5+ years in technical delivery, quality, or program management
  • Hands-on experience delivering data or evaluation work to AI labs or enterprise ML teams
  • Working fluency with how frontier models are evaluated
  • Proven people/vendor leadership

Responsibilities

  • Study public benchmarks and eval targets, and turn them into proposals and sample packages
  • Design and build the sample packages, working with subject-matter experts
  • Recruit, brief, calibrate, and review a pool of experts across coding, agentic/tool-use, and STEM/reasoning
  • Own pilots end to end: scoping, SOW, staffing, production, QC, and delivery

Key facts

Other skills

  • Quality Assurance
  • Communication
  • Collaboration
  • Problem Solving

About the company

Anyone AI logo

Anyone AI

E-Learning / EdTech

We invest in talent from Latam to bridge the talent gap in AI. Join our AI community: www.anyoneai.com We are AI / ML experts and second-time entrepreneurs, members of the founding team at Deep Vision AI (acquired in early 2020). We've worked with many Fortune 500 companies completing multiple projects in the early days of AI while leveraging remote talent from LatAm. We are VC-backed from day 1 by top global investors like GFC -Global Founders Capital- (investors in Facebook, LinkedIn, Slack, Canva, Trivago, etc), Canvas Ventures (early investor in Coursera), Latitud Fund (the largest community of angel investors for LatAm including investments in QuintoAndar, La Haus, Clara, Platzi, OnTop, Pomelo, etc), among other investors.

Company details

IndustryE-Learning / EdTech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Reports to: CEO

Owns: data proposals, sample development, quality, and pilot delivery

Location: Remote / Latam / US


The role

You will own Anyone AI’s data initiatives and proposals to AI labs, from the data proposal or responding to requests, through pilot delivery. You own how we build proposals and develop the sample packages and benchmarks: frontier-grade packages across reasoning, coding, agents, and tool use, multi-modal and others, produced in collaboration with subject-matter experts, with expert-verified ground truth, multi-model headroom results, and QC that survives buyer-side scrutiny. You are the person who designs the sample that demonstrates our quality, converts pilots into production engagements. On a small team, this is the operational center of the Human Data Division.

Responsibilities

  • Proposals & requests. Study public benchmarks and eval targets, and turn them into proposals and sample packages that demonstrate capability and win the work. Respond to lab data requests and pilots.

  • Sample & benchmark development. Design and build the sample packages, working with subject-matter experts. Every package meets the bar of our current sample set:

    • Expert-verified, exact-match-checkable ground truth and gold reasoning trajectories.

    • Multi-model evaluation showing real headroom, and proof the task discriminates the model, not just that it's hard.

    • Rigorous QC structure: calibration layers, severity-weighted rubrics, deterministic verifiers, evidence maps, etc.

  • Subject-matter experts. Recruit, brief, calibrate, and review a pool of experts across coding, agentic/tool-use, and STEM/reasoning. Raise their output to our standard and keep it there; be the arbiter of what "correct" and "frontier-difficulty" mean.

  • Lab relationships. Be a direct point of contact for lab partners on Slack and calls, with support from the CEO and the wider team. Keep senior lab contacts informed, surface what they actually need, and pull in the CEO and subject-matter experts when the conversation calls for it.

  • Pilot delivery. Own pilots end to end: scoping, SOW, staffing, production, QC, and delivery. Nothing ships before it's lab-ready, and nothing comes back rejected as "not frontier-level" without us already knowing why.

Experience

  • Originated data or benchmark proposals for AI labs, translated eval targets into sample tasks that demonstrate capability, and owned the engagement through delivery.

  • Deep evaluation and quality expertise: LLM benchmarking, with real strength in code-model evaluation.

  • Built QC processes and artifact standards that met enterprise or lab requirements, and set a quality bar a team of experts was held to.

  • Thrives in ambiguous, fast-moving environments where the rules are still being written, and delivers under pressure.

Qualifications

  • 5+ years in technical delivery, quality, or program management, with recent experience in AI/ML data, model evaluation, or benchmarking.

  • Hands-on experience delivering data or evaluation work to AI labs or enterprise ML teams, scoping through delivery.

  • Working fluency with how frontier models are evaluated: benchmarks, rubrics, pass rates, headroom, and what makes a task discriminate a model.

  • Proven people/vendor leadership, you've recruited, calibrated, and held a team or expert pool to a quality standard.

  • Fluent English. Spanish is a nice to have.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Anyone AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.