Logo for Cantina

Machine Learning Enginer, Core Evaluations

Role overview

Qualifications

  • Strong experience designing metrics that capture model performance
  • Strong experience with designing user studies on Mechanical Turk or similar platforms
  • Strong experience with model training and fine-tuning for model evaluation
  • Very strong engineering and programming skills

Responsibilities

  • Designing model evaluation pipelines for models in development and production
  • Designing user studies for subjective model evaluations
  • Converting requirements into measurable metrics
  • Designing and developing automated evaluation dashboards to monitor model performance and compare results

About the company

Cantina logo

Cantina

IT Services & IT Consulting

Building the first social AI platform

Company details

Company typeScaleup
IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are seeking an experienced Machine Learning Engineer (MLE) to focus on audio model evaluation, specifically for speech generation and recognition models.

This role involves designing and developing comprehensive model evaluation pipelines for both development and production environments, as well as creating automated dashboards for reporting evaluation results.

As the founding member of our evaluation team, the ideal candidate is expected to leverage their experience to lead our evaluation efforts and play a key role in the future growth of the evaluation team.

What You’ll Do:

  • Designing model evaluation pipelines for models in development and production

  • Designing user studies for subjective model evaluations.

  • Converting requirements into measurable metrics.

  • Designing and developing automated evaluation dashboard to see model performances and compare results.

  • Training new models to capture new and different evaluation metrics.

  • Communicating with the model team to help design better models based on the evaluation results.

  • Communicating with the data team to help decide the type of data necessary to improve model performance.

  • Communication with the product-manager to make sure product requirements are correctly measured.

  • Help grow the evaluation team as the founding member.

  • Lead the evaluation team in the future.

What You’ll Bring:

  • Strong experience and intuition for designing metrics that capture model performance.

  • Strong experience with designing user studies on Mechanical Turk or similar platforms. .

  • Strong experience with model training and fine-tuning for model evaluation.

  • Strong statistical knowledge and experience to statistically compare evaluation results and take decisions.

  • Very strong engineering and programming skills.

  • Experience with training ASR, TTS models.

  • Experience at ML teams working on large-scale machine learning problems. (>3B models with >1m hours of data)

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.