Logo for Clera

Data Scientist — Agent Evaluations & Quality

Role overview

Qualifications

  • 5+ years in data science, machine learning, or analytics roles
  • Demonstrated experience designing and implementing evaluation frameworks
  • Production-quality Python and SQL skills
  • Strong evaluation methodology chops

Responsibilities

  • Architect and maintain automated evaluation pipelines for agent quality
  • Translate complex agent behaviors into success criteria
  • Build gold datasets and regression suites for workflows
  • Run rigorous offline experiments and leverage production evidence

Key facts

  • Remote from: California (USA)
  • Full time
  • Senior (5-10 years)
  • Data Scientist
  • English

Hard skills

Other skills

  • Analytical Skills

About the company

Clera logo

Clera

Staffing & Recruiting

Clera is the first AI talent agent: a personal headhunter that acts on behalf of top talent. We connect you to your dream job.

Company details

IndustryStaffing & Recruiting
Company size1 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the Role

This is an applied data science role focused on measuring, understanding, and improving the quality of AI agents that handle real-world tasks — scheduling, email, browser automation, business software, and more. You'll sit at the intersection of evaluation design, statistics, and production systems, building the feedback loops that drive engineering and product decisions. The work matters because agent quality is hard to measure and easy to get wrong.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate complex, multi-step agent behaviors into explicit success criteria — including pass, partial-pass, and failure definitions.

  • Build gold datasets and regression suites covering common workflows, edge cases, and adversarial scenarios.

  • Define and track metrics spanning task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.

  • Analyze traces, tool calls, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Run rigorous offline experiments and leverage production evidence to compare models, prompts, and capability implementations.

  • Build dashboards and reports that make evaluation results clear and actionable for engineering, product, and leadership.

  • Partner with engineers to recommend improvements and verify fixes raise quality without unacceptable regressions.

What We're Looking For

  • 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks and grading systems for ML or AI products in production.

  • Production-quality Python and SQL skills; ability to build automated pipelines and conduct analysis at scale.

  • Strong evaluation methodology chops: success criteria definition, dataset construction, metric selection, and spotting misleading benchmarks.

  • Solid statistical and experimental design knowledge — sampling, variance, uncertainty quantification, bias, confounding, and significance testing for non-deterministic systems.

  • Experience developing ground-truth data: labeling guidelines, annotation QC, ambiguity resolution, and dataset maintenance.

  • Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes.

  • Ability to connect quantitative patterns to individual traces and identify failure origins across model, prompt, context, tools, and application logic.

  • Clear communicator who can convey evaluation results, methodology, and trade-offs to both technical and non-technical stakeholders.

  • Experience with agentic systems, multi-step task evaluation, or consumer-facing production ML is a strong plus.

Location

On-site in Palo Alto, California, United States. Visa sponsorship is not available.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Scientist Related jobs

Other jobs at Clera

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.