Logo for Samsung Food

AI Quality & Evaluation Lead

Role overview

Qualifications

  • Experience running end-to-end evaluation loop on conversational or generated-text output
  • Background in conversation design, AI product quality, model behavior/policy, content design, or UX research
  • Strong qualitative coding skills

Responsibilities

  • Build failure taxonomy and golden set for AI coaching experience
  • Develop a rubric with binary pass/fail criteria for evaluation
  • Validate an LLM judge against human-generated grades
  • Create a written playbook to guide future evaluations and handover

Key facts

Hard skills

Other skills

  • Communication
  • Problem Solving

About the company

Samsung Food logo

Samsung Food

Computer Software / SaaS

Samsung Food powers the creation, discovery, personalization, and monetization of food content online, in-store, and at home. Samsung Food partners with the world’s largest retailers, publishers, CPG brands, and health companies to help them connect with their consumers and drive product purchases at every stage of their food journey - from inspiration and consideration to purchase and preparation. Samsung Food, previously known as Whisk, was founded in 2012 and was acquired by Samsung Next in March 2019. Samsung Food was certified as a Most Loved Workplace® after rigorous research and analysis conducted by the Best Practice Institute (BPI).

Company details

IndustryComputer Software / SaaS
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Samsung Food (Whisk)

At Samsung Food (you might know us as Whisk), we are pioneering the future of everyday care and health — helping individuals and families around the world live better through connected, proactive, and personalized experiences. Our mission is to connect food, health, and home across the Samsung ecosystem, turning data from your devices into personalized guidance that supports you and your loved ones.

We work at the intersection of AI, nutrition, behavior change, and digital health. The Samsung Food app was included in Google Play’s Best of 2020 Everyday Essentials list, has been regularly featured on the Apple App Store, and was nominated for a 2021 Webby Award. We’ve also launched Vision AI, which recognises food and meals using our global Food Genome, and developed deep integrations with Samsung Health and SmartThings.

We are a remote-first team of more than 100 people in over 30 countries, united by our commitment to innovation and impact. As part of Samsung Electronics, we combine the agility of a startup with the scale of the world’s largest consumer electronics company. Join us and help build technology that improves daily life for tens of millions of people across phones, wearables, TVs, and kitchen appliances in the years to come.

Learn more about how we’ve shaped a high-performing global team over the past 14 years.

Summary

Overview

We're building an AI coaching experience inside Samsung Health. Every piece of coaching a user sees is generated, not written. Our platform partner owns the evaluation runtime and functional evals (did the pipeline execute, did the tool call succeed). Nobody owns the layer above that: is the output actually good — right framing, right tone, right structure, does it deliver the coaching logic, does it feel like it's talking to me.

This engagement exists to build that layer: author the failure taxonomy, codify it into a binary rubric, validate an LLM judge against your own human grades, and hand the whole thing to an internal owner. We have a v0.1 rubric and golden scenarios from our product lead, and a functional rubric owned by engineering. We don't have the experience rubric — that's what this engagement produces.

Deliverables by Phase (Overall project ~13 weeks)

1. Failure taxonomy and golden set (within 4 weeks from the contract start date)

At least 150 traces hand-graded with open-coded notes. At least 40 of them come from sparse-data synthetic profiles. The taxonomy has 5–10 categories, each with a frequency count, and the top three failures are shown with their share of all failures. Saturation is evidenced: the last 20 traces produced no new category. At least 50 traces are frozen as a holdout. Head of Product signs off the taxonomy.

2. Rubric v1 and calibration (within 4 weeks from phase 1 completion)

One binary pass/fail criterion per taxonomy category, each with a definition, a pass example and a fail example. Written annotation guidelines. At least 30 traces are double-coded with a second grader, and every disagreement is logged and resolved in writing. A guideline revision log records each change and its reason. Head of Product signs off the rubric.

3. Validated judge and weekly readout (within 4 weeks from phase 2 completion)

Judge prompts for every criterion. A validation report giving true positive rate and true negative rate per criterion on both the dev set and the holdout. A weekly readout on helpfulness, relevance and tone, run at least twice: once by the consultant, and once by the internal owner with the consultant shadowing. A documented revalidation routine, triggered by every model or prompt change and run quarterly regardless.

4. Playbook and handover test (within 1 week from phase 3 completion)

A written playbook covering the whole method: grading, updating the taxonomy, revising the rubric, revalidating the judge and running the readout. The handover test passes when the internal owner reruns the judge validation unaided, gets within ±5 points, and runs one weekly readout with no consultant involvement.

What This Role Is Not

- Not building eval infrastructure, harnesses, or dashboards (owned elsewhere)

- Not functional/module-level evals or accuracy metrics

- Not clinical safety validation

- Not high-volume annotation — grading a few hundred traces to build the taxonomy; volume work goes to vendors/the judge once the rubric exists

What We're Looking For

You've run this loop end-to-end at least once on conversational or generated-text output: error analysis on real traces → a failure taxonomy you built yourself → binary criteria and annotation guidelines → an LLM judge validated against your own labels.

Background: conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, or RLHF. Applied linguistics and product management backgrounds also fit.

You'll be strong in this role if you:

- Can describe a failure mode you personally discovered by reading output — one nobody had named before you

- Are comfortable being the arbiter and documenting what you overruled and why

- Understand the rubric is discovered through grading, not written upfront

- Can explain why a judge that agrees with you 94% of the time might still be useless

- Write annotation guidelines that don't need interpreting

- Are comfortable working in notebooks and spreadsheets (no production code required)

Not required: domain expertise in nutrition, weight management, or behaviour change — we have that in-house.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Samsung Food

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.