Logo for Workera

Senior Data Scientist, AI Scoring & Evaluation

Role overview

Qualifications

  • 4+ years of production data project experience in an AI startup or fast-moving product environment
  • Proficient in Python and SQL
  • Strong statistical and analytical skills
  • Experience with LLM-based systems in production

Responsibilities

  • Own scoring quality end to end: evaluator design, rubric anchoring, calibration
  • Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection
  • Define and publish assessment quality KPIs
  • Lead analytics investigations behind assessment decisions

About the company

Workera logo

Workera

IT Services & IT Consulting

Workera.ai is the precision upskilling platform helping enterprises, governments and individuals upskill and reskill to meet the critical demand for key technological capabilities including data science, machine learning, and artificial intelligence. The Workera platform provides AI-driven mentorship at scale with its adaptive assessments and personalized learning plans that drive measurable results to close the skills gap.

Company details

Company typeScaleup
IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Senior Data Scientist - Assessment Scoring & Evaluation

Everyone's racing to build AI. Workera exists for the 8 billion people who work alongside it.

While the world's attention is on creating new tools, someone has to solve the other side of the equation: the humans. The workforce is going through the biggest transformation in a generation and most organizations are navigating it blind, without the data to understand what their people can actually do, where the gaps are, or how to close them fast enough.

As Workera moves into higher-stakes decisions, you're the person who can prove a measurement is accurate, fair, and defensible. You'll own the measurement system that turns signals into the skill data our customers act on: the evaluation harnesses, the calibration methods, the quality monitoring behind every score. You'll also shape how our AI software itself is designed, from how components are orchestrated and prompted to the evaluation frameworks that decide whether they're good enough to ship.

You're not just running analyses, you're building the trust layer underneath a product that increasingly makes high-stakes calls about people's careers. If you want a role where your judgment becomes the thing customers, auditors, and your own engineering team rely on, and where the measurement problems are still being invented rather than optimized, this is it.

Workera's skills intelligence platform is critical infrastructure for the AI era: the layer that lets organizations understand, mobilize, manage, and develop their talent with precision. We're trusted by the Fortune 500, powered by proprietary AI agents, and built by a small, senior team, which means what you ship here has outsized reach.

WHY THIS ROLE EXISTS

Workera is scaling fast: more customers, more use cases, more fields evaluated, higher stakes on every measurement. That scale changes what quality means. What we once crafted and inspected by hand now needs monitoring that catches issues before customers do, and remediation that resolves them fast.

This role sits at the intersection of Engineering, Assessment Science, and Design, partnering closely with the product engineers who build our scoring pipeline. You're the data layer that pieces those disciplines together: the foundation the rest of our scoring is built on, and the reason a construct definition from Assessment Science actually turns into a number a customer can trust.

YOUR TEAM

You report to Dr. Taylor Sullivan, Workera’s VP of Product and Assessments, and work day to day with two groups: The assessment science team decides what we're measuring and how a skill turns into test content, they own construct definition (what "good at X" actually means), blueprinting (how an assessment is structured), and content authoring (writing the actual questions). You partner with them on how scoring methods serve that intent. The assessment tech team builds and maintains the scoring pipeline; you work embedded with them: writing specs, implementing and coordinating improvements, and validating that what ships meets the quality bar. You also work with data engineering on data availability. It's a small, cross-functional team, and you sit inside product decisions rather than in a separate research function.

WHAT YOU'LL OWN

Build trust by owning our scoring mechanism. These are the outcomes you're accountable for:

  • Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability for accuracy against expert benchmarks
  • Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and the pre-release gate that every scoring change must clear
  • Define and publish assessment quality KPIs (human to AI agreement, reliability, classification accuracy, bias indicators, latency, cost per assessment) on dashboards anyone in the company can reference
  • Lead the analytics investigations behind assessment decisions: performance studies, impact simulations, root cause analysis when scores look wrong
  • Ship measurement improvements end to end: implement and coordinate with tech, own the spec, the validation, and the quality bar in both cases
  • Be Workera's liaison between AI governance and InfoSec: stay current on frameworks like GDPR and the EU AI Act, and turn that into the evidence we show in audits, security reviews, and enterprise diligence when customers ask how we use AI in our assessments
  • Empower partners across Product, Engineering, and GTM to get the data and analysis they need autonomously, by building AI tooling and documentation rather than answering each request yourself

HOW YOU'LL RAMP

We don't expect you to figure it out alone. Here's what great looks like at each stage:

First 30 Days: Learn the Machine

  • Immerse yourself in Workera's platform, customers, and the problems we're solving. You'll shadow key workflows and understand how AI is embedded in day-to-day operations across teams
  • Deliver at least one analytics investigation that changes a product decision
  • Sit in on an enterprise or auditor conversation about score quality to see what evidence customers actually ask for

By 60 Days: Ship Something Real

  • Own your first meaningful deliverable and demonstrate end-to-end execution
  • Troubleshoot scoring end to end independently
  • Ship one measurable scoring improvement, validated against expert labelled data, through the pre-release gate

By 90 Days: Multiply Your Impact

  • Operate with full autonomy in your domain; your team relies on your judgment
  • Have built or deployed at least one AI-assisted workflow that the team adopts
  • Set the scoring roadmap: identify the gaps, prioritise what gets built, and drive those initiatives to completion with Engineering
  • Build systems that improve themselves: evaluation loops that flag their own drift, calibration that updates on new expert labels, tooling that gets better as it is used

 

We're a fast-moving company -- the scope and shape of this role will evolve as we do.

 

WHAT YOU BRING

We're looking for signal, not checkboxes. In rough order of what matters most:

  • You've driven production data science work independently.
  • 4+ years (or 3+ with demonstrated end-to-end ownership) of a production data project ideally in an AI startup or fast-moving product environment where you scoped investigations yourself rather than picking up defined tickets.
  • Python and SQL are assumed; you shouldn't need an engineer to run an analysis.
  • Real statistical and analytical depth. You can turn, "is our scoring reliable?" into a defined study with a defensible method and a clear answer fast, and without being handed the design.
  • You can read this kind of data. Quantitative work with educational, learning, or assessment data or a clear appetite to go deep on it fast. We're looking for someone who can interpret our assessment data fluently and wants to become expert in how it's built.
  • You've worked with LLM-based systems in production. Prompt design, evaluating model output against human judgment, and a real sense of where these systems fail.
  • You explain quantitative results to non-technical audiences, including executives and customers, without losing the nuance.
  • A plus, not a requirement: background in psychometrics or educational measurement. Our assessment scientists will coach you β€” but you should want to learn it.

 

HOW WE WORK: AI IS THE DEFAULT

At Workera, AI isn't a feature we sell -- it's how we operate. Every team member is expected to:

  • Use AI daily. AI assistants, copilots, and automation tools are part of your stack -- not optional extras. We expect you to actively experiment with new tools and push the boundary of what's possible in your function.
  • Build your own leverage. Our marketers write code. Our PMs build automations. Our ops team deploys agents. If a workflow can be automated, you're expected to automate it.
  • Think in systems, not tasks. We value people who build repeatable, scalable solutions over people who grind through one-off work. Your goal is to make your function run smarter, not just harder.

 

AI fluency is a cultural expectation, not a line item on a job description.

 

ABOUT WORKERA

We're a Silicon Valley company backed by NEA, Jump Capital,  and Owl Ventures. Our founder is Kian Katanforoosh, an award-winning Stanford Computer Science Lecturer who has taught AI to over 1 million people. Our Chairman is Dr. Andrew Ng,  co-founder of Coursera, CEO of DeepLearning.AI, and founding lead of the Google Brain project.

Our clients include Accenture, Siemens Energy, Samsung, and the United States Air Force.

Named to Fast Company's Most Innovative Companies list alongside Microsoft and Canva. Recognized by the World Economic Forum's Tech Pioneers, Inc 5000, and Josh Bersin's HR Tech AI Trailblazers. In a world where every company claims to 'do AI',  at Workera, it's actually in our DNA.

We're learners, builders, and dreamers. Join us.

Workera is committed to providing an inclusive and respectful environment where equal employment opportunities are available to all applicants and employees. We do not discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law. Hiring decisions are based on qualifications, merit, mindset, and business need.

About Workera
We're a Silicon Valley company backed by NEA, Jump Capital,  and Owl Ventures. Our founder is Kian Katanforoosh, an award-winning Stanford Computer Science Lecturer who has taught AI to over 1 million people. Our Chairman is Dr. Andrew Ng,  co-founder of Coursera, CEO of DeepLearning.AI, and founding lead of the Google Brain project.

Our clients include Accenture, Siemens Energy, Samsung, and the United States Air Force.

Named to Fast Company's Most Innovative Companies list alongside Microsoft and Canva. Recognized by the World Economic Forum's Tech Pioneers, Inc 5000, and Josh Bersin's HR Tech AI Trailblazers. In a world where every company claims to 'do AI',  at Workera, it's actually in our DNA.

We're learners, builders, and dreamers. Join us.
Workera is committed to providing an inclusive and respectful environment where equal employment opportunities are available to all applicants and employees. We do not discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law. Hiring decisions are based on qualifications, merit, mindset, and business need.



Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Data Scientist Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.