Logo for Featherless AI

AI Researcher — Distillation

Role overview

Qualifications

  • Strong background in machine learning research
  • Hands-on experience with model distillation or closely related topics (compression, pruning, quantization, representation learning)
  • Publication experience (conference or journal papers, workshop papers, or arXiv preprints)
  • Fluency in PyTorch (or equivalent) and research-grade experimentation

Responsibilities

  • Design and evaluate model distillation techniques (teacher-student training, self-distillation, layer-wise distillation, representation matching)
  • Research tradeoffs between model size, latency, memory, and accuracy; develop novel distillation approaches for large language models, long-context or specialized architectures
  • Run large-scale experiments and ablations; analyze results rigorously; collaborate with engineers to productionize research outcomes
  • Publish research findings and contribute to internal notes, technical blogs, and open-source projects when appropriate

About the company

Featherless AI logo

Featherless AI

Artificial Intelligence & Machine Learning Services

We enable serverless inference via our GPU orchestration and model load-balancing system. We unlock fine-tuning by enabling organizations to size their server fleet to throughput needs, not number of models in the catalogue. See it in action on our public cloud, which offers inference for 4,200+ open weight models.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size1 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the Role

We’re looking for an AI Researcher focused on model distillation to help us push the frontier of efficient, high-performance models. You’ll work on turning large, expensive models into smaller, faster, and more deployable systems—while maintaining or improving quality.

This role is ideal for someone who enjoys publishing research, working close to real systems, and seeing their ideas move from papers → code → production.

What You’ll Work On

  • Design and evaluate model distillation techniques (teacher–student training, self-distillation, layer-wise distillation, representation matching, etc.)

  • Research tradeoffs between model size, latency, memory, and accuracy

  • Develop novel distillation approaches for:

    • Large language models

    • Long-context or specialized architectures

    • Inference-constrained environments

  • Run large-scale experiments and ablations; analyze results rigorously

  • Collaborate with engineers to productionize research outcomes

  • Write and submit research papers to top-tier venues (NeurIPS, ICML, ICLR, COLM, etc.)

  • Contribute to internal research notes, technical blogs, and open-source projects when appropriate

What We’re Looking For

Required

  • Strong background in machine learning research

  • Hands-on experience with model distillation or closely related topics (compression, pruning, quantization, representation learning)

  • Publication experience (conference or journal papers, workshop papers, or arXiv preprints)

  • Solid understanding of deep learning fundamentals (optimization, training dynamics, generalization)

  • Fluency in PyTorch (or equivalent) and research-grade experimentation

  • Ability to clearly communicate research ideas, results, and limitations

Nice to Have

  • Experience distilling large language models

  • Work on efficiency-focused research (latency, memory, throughput)

  • Experience with long-context models or non-Transformer architectures

  • Open-source contributions in ML or research tooling

  • Prior startup or applied research experience

Why Join Us

  • Real ownership over research direction at a Series A stage

  • Strong support for publishing and open research

  • Tight feedback loop between research and real-world deployment

  • Access to meaningful compute and production-scale problems

  • Small, highly technical team with deep ML and systems expertise

Example Backgrounds

  • ML researchers from academia transitioning to industry

  • Research engineers with published work in model efficiency

  • PhD / Post-doc graduates or industry researchers who still want to publish

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

AI Specialist Related jobs

Other jobs at Featherless AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.