Logo for Protege

Applied Healthcare Researcher

Role overview

Qualifications

  • Advanced degree (PhD or Master's plus 3+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field.
  • Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data.
  • Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar.
  • Strong Python and SQL, with the ability to work independently against large datasets.

Responsibilities

  • Serve as the primary technical and research point of contact for healthcare customer conversations.
  • Translate a lab's model-development goals into concrete, feasible data strategies.
  • Develop and evaluate methods to demonstrate that a dataset can support a customer's training or evaluation objective.
  • Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data.

Key facts

Other skills

  • Collaboration
  • Communication
  • Problem Solving

About the company

Protege logo

Protege

Computer Software / SaaS

The biggest unmet need in AI today is getting access to the right training data. Data holders often don’t know where to start and are rightly concerned about governance, intellectual property, and security implications. AI companies can spend years finding and negotiating access to the data they need. Protege is solving these problems by providing an easy-to-use platform to connect data holders with vetted data users.

Company details

IndustryComputer Software / SaaS
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Company Overview:

We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.

Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.

We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.

Role Overview

We are hiring Applied Healthcare Researchers to join a team within DataLab focused entirely on healthcare training data.

Our customers are researchers at the frontier labs and AI startups building specialized healthcare models. They come to us with model-development problems, not dataset specifications. Figuring out which healthcare data actually solves their problem, and proving that it does, is the research question we answer in DataLab.

In this role you will work directly with researchers at those labs to understand what they're trying to train or evaluate, determine what healthcare data can support it, and do the research needed to demonstrate that it will. This is fast-iterating, customer-facing research on a customer's timeline. You will be the primary technical and research link to the customer — not a technical resource brought in for credibility, but the person driving the conversation and pulling in the solutions, engineering, and data partnerships teams as needed.

Core Responsibilities

Customer Research Partnership

You will be the research partner to AI researchers at frontier labs and startups who are working on healthcare problems.

• Serve as the primary technical and research point of contact for healthcare customer conversations.

• Translate a lab's model-development goals into concrete, feasible data strategies.

• Help customers scope opportunities and identify the highest-value data available to them.

• Explain data limitations, tradeoffs, and potential biases to technically sophisticated stakeholders while grounding conversations in what real-world data actually looks like.

• After delivery, answer the research questions customers raise about the data we provided. Delivery is not the end of the relationship.

Applied Research & Method Development

Curating the right data product is a research problem, and you'll own solving it.

• Develop and evaluate methods — fine-tuning, LLM-based extraction, classification, rules-based approaches, or whatever the problem calls for — to demonstrate that a dataset can support a customer's training or evaluation objective.

• Design and run feasibility research pre-contract: can this data support this model objective, at what quality, with what caveats.

• Build the evidence base that makes a data strategy credible — benchmarks, validation analyses, error characterization, and honest assessments of where the data falls short.

• Partner with the Assessments team on healthcare benchmarks across modalities.

Data Feasibility & Dataset Strategy

• Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data.

• Identify proxy variables or alternative dataset structures when the ideal variable doesn't exist.

• Analyze partner and source datasets — schema, field availability, quality, completeness, and required transformations.

• Contribute to our point of view on which healthcare data matters most for which modality and which stage of model development.

• Help evaluate new data partners and identify datasets worth acquiring before a customer asks for them.

Reusable Research & Scaling

• Produce reusable research, evidence, and technical collateral rather than starting from scratch for each opportunity.

• Identify where a successful one-off approach should become a repeatable workflow, and work with Product and Engineering to operationalize it.

• Help expand proven healthcare datasets across multiple customers instead of selling them once.

Cross-Functional Collaboration

• Work with Solutions and FDEs from the beginning of an opportunity.

• Coordinate with Healthcare Data Partnerships on sourcing and with Product and Engineering on tooling.

Required Experience & Skills

• Advanced degree (PhD or Master's plus 3+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience.

• Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data.

• Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training.

• Strong Python and SQL, with the ability to work independently against large datasets.

• Experience designing evaluations — measuring data quality and dataset representativeness.

• Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans.

• Comfort operating on a customer's timeline without lowering the standard of the research.

Ideal Profile

The ideal candidate:

• Is energized by working directly with customers, and specifically by working with other researchers as peers.

• Moves fast on messy, real-world problems and knows which corners can and cannot be cut.

• Is rigorous about what the data can and cannot support, and willing to tell a customer when the answer is no.

• Enjoys the full arc — scoping a vague problem, doing the research, and showing the result to the person who asked for it.

• Thinks about leverage: builds the reusable version rather than the one-off when it's worth doing.

About DataLab

DataLab exists because truly useful data is rare — and the frontier of AI development only moves forward when high-quality data makes it possible.

We believe data is one of the most underdeveloped layers of the AI stack. Our work focuses on building and evaluating high-value datasets grounded in real-world workflows and economically meaningful tasks. Our research spans data quality, evaluation design, privacy-preserving transformation, and task-grounded AI training data.

Protege Values

Pass the Loved Ones’ Test

We act with integrity and do the right thing — especially when it’s hard and no one is watching.

Always Find a Way

We are resourceful, resilient builders who solve hard problems and push through obstacles.

Go Fast and Grow Fast

Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.

Practice Kindness and Candor

We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.

Deliver Together

We win as one team. Collaboration, accountability, and shared ownership drive our success.

Own the Outcome. Hone the Craft.

We take pride in our work, sweat the details, and continuously raise the bar for excellence.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Researcher Related jobs

Other jobs at Protege

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.