Logo for Innodata Inc.

Research Scientist, Speech & Audio

Role overview

Qualifications

  • Roughly 5+ years of hands-on industry experience in speech or audio ML.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree preferred.
  • Trained and evaluated speech or audio models with strong PyTorch fundamentals.
  • Fluency in toolchains and metrics for speech work such as ESPnet, NeMo, SpeechBrain, or Kaldi.

Responsibilities

  • Define how Innodata designs, structures, and evaluates audio data for speech and audio models.
  • Build evaluation methodology that goes beyond word error rate and assess speech system performance.
  • Decide how audio data should be structured and enriched for model objectives.
  • Partner with teams to ensure good data and evaluation specifications are met.

Key facts

Hard skills

Other skills

  • Communication

About the company

Innodata Inc. logo

Innodata Inc.

Artificial Intelligence & Machine Learning Services

(NASDAQ: INOD) Innodata is a global data engineering company delivering the promise of AI to many of the world’s most prestigious companies. We provide AI-enabled software platforms and managed services for AI data collection/annotation, AI digital transformation, and industry-specific business processes. Our low-code Innodata AI technology platform is at the core of our offerings. In every relationship, we honor our 30+ year legacy delivering the highest quality data and outstanding service to our customers.

Company details

Company typeLarge
IndustryArtificial Intelligence & Machine Learning Services
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

Where models actually differ now is robustness across accents, noise, and code-switching; speaker diarization; the naturalness of generated speech; and latency under streaming. Measuring those honestly, and building the data that trains for them, is gated as much by data and evaluation design as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing speech and audio models, and we are hiring a Research Scientist to own the science behind it. 

You will partner directly with the customers and frontier labs building ASR, text-to-speech, speech-to-speech and conversational voice, diarization, and audio-language models, as interested in the data behind them as in the models themselves. Your work is judgment: which conditions and languages a benchmark must cover to be honest, what a transcription convention should be for a given objective, and when an automated metric can be trusted versus when a human ear is required. You will also partner closely with our transcription and linguistics lead, whose standards directly shape what the models learn. 

What You’ll Own:

  • You will define how Innodata designs, structures, and evaluates audio data for speech and audio models, and you will validate those choices experimentally. Concretely, you will: 
  • Translate the requirements of speech and audio models — ASR, text-to-speech and speech generation, speech-to-speech and conversational voice, speaker diarization and verification, audio-language models, and streaming systems — into concrete data specifications: modalities, transcription and annotation schemas, sampling, and evaluation criteria. 
  • Build evaluation methodology that goes past word error rate — semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency — and know when automated metrics hold and when they don't. 
  • Decide how existing and incoming audio should be structured, enriched, and sampled for coverage that fits the model objective — across languages, accents, and acoustic conditions (studio, real-world, telephonic), speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech. 
  • Partner with the transcription and linguistics lead to turn model objectives into transcription specifications, and to quantify how transcription conventions and quality move ASR and speech-model results. 
  • Partner with the audio solutions and engineering team so the audio we collect is built for the model objective: you specify what good data and evaluation require, and they scope programs with customers and capture audio to spec. 
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement. 
  • Design adversarial and stumping evaluations — noisy, accented, and adversarial audio — that surface where speech systems fail, and turn those failures into better data. 
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with. 
  • Work with annotation teams, subject-matter experts, and the synthetic- and augmented-audio pipeline to turn specifications into operational plans. 

You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in speech or audio ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end. 
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred. 
  • Trained and evaluated speech or audio models yourself — ASR, TTS, speech-to-speech, speaker, or audio-language models — with strong PyTorch fundamentals. 
  • Fluency in the toolchains and metrics speech work runs on: ESPnet, NeMo, SpeechBrain, or Kaldi, HuggingFace, forced alignment, and WER/CER and the metrics that go beyond them. 
  • Hands-on experience with multilingual, accented, dialectal, low-resource, or code-switched speech, and with synthetic or augmented audio (TTS pipelines, noise and room-response simulation). 
  • A way of thinking in datasets: you have built evaluation sets, reasoned about coverage across conditions, and argued about what makes speech data good for a given objective. 
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS. 
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation. 
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — for example spoofing and voice-cloning robustness, or bias across accents and languages. 

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Research Scientist Related jobs

Other jobs at Innodata Inc.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.