Clera
Staffing & Recruiting
See how your profile stacks up against this role.
We compared the job requirements to your profile to show where you're strong and where you fall short.
This is an applied data science role focused on measuring, understanding, and improving the quality of AI agents that handle real-world tasks — scheduling, email, browser automation, business software, and more. You'll sit at the intersection of evaluation design, statistics, and production systems, building the feedback loops that drive engineering and product decisions. The work matters because agent quality is hard to measure and easy to get wrong.
Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.
Translate complex, multi-step agent behaviors into explicit success criteria — including pass, partial-pass, and failure definitions.
Build gold datasets and regression suites covering common workflows, edge cases, and adversarial scenarios.
Define and track metrics spanning task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
Design deterministic and model-based graders; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.
Analyze traces, tool calls, and production outcomes to identify root causes and build a useful failure taxonomy.
Run rigorous offline experiments and leverage production evidence to compare models, prompts, and capability implementations.
Build dashboards and reports that make evaluation results clear and actionable for engineering, product, and leadership.
Partner with engineers to recommend improvements and verify fixes raise quality without unacceptable regressions.
5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.
Demonstrated experience designing and implementing evaluation frameworks and grading systems for ML or AI products in production.
Production-quality Python and SQL skills; ability to build automated pipelines and conduct analysis at scale.
Strong evaluation methodology chops: success criteria definition, dataset construction, metric selection, and spotting misleading benchmarks.
Solid statistical and experimental design knowledge — sampling, variance, uncertainty quantification, bias, confounding, and significance testing for non-deterministic systems.
Experience developing ground-truth data: labeling guidelines, annotation QC, ambiguity resolution, and dataset maintenance.
Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes.
Ability to connect quantitative patterns to individual traces and identify failure origins across model, prompt, context, tools, and application logic.
Clear communicator who can convey evaluation results, methodology, and trade-offs to both technical and non-technical stakeholders.
Experience with agentic systems, multi-step task evaluation, or consumer-facing production ML is a strong plus.
On-site in Palo Alto, California, United States. Visa sponsorship is not available.
After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.
Marcus Rivera
Chief Revenue Officer

Aspire General Insurance

John Deere

Clera Inc.

EPIC Insurance Brokers & Consultants

Wealthfront

Clera

Clera

Clera