Logo for Innodata Inc.

Principal Speech Data Linguist

Role overview

Qualifications

  • Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality
  • A Bachelor's degree in linguistics, phonetics, or computational linguistics, or a closely related field
  • Fluency in phonetic transcription and IPA, plus hands-on experience with acoustic and phonetic analysis of speech
  • Deep experience with audio segmentation and its conventions

Responsibilities

  • Define transcription and segmentation standards, style guides, and annotation conventions
  • Establish and run the quality frameworks behind segmentation and transcription work
  • Own the quality lifecycle for transcription and segmentation deliverables end to end
  • Design the human-in-the-loop workflows for improved ASR quality

Key facts

Hard skills

Other skills

  • Quality Assurance
  • Communication
  • Mentorship

About the company

Innodata Inc. logo

Innodata Inc.

Artificial Intelligence & Machine Learning Services

(NASDAQ: INOD) Innodata is a global data engineering company delivering the promise of AI to many of the world’s most prestigious companies. We provide AI-enabled software platforms and managed services for AI data collection/annotation, AI digital transformation, and industry-specific business processes. Our low-code Innodata AI technology platform is at the core of our offerings. In every relationship, we honor our 30+ year legacy delivering the highest quality data and outstanding service to our customers.

Company details

Company typeLarge
IndustryArtificial Intelligence & Machine Learning Services
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years. 

This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself. 

What You’ll Own:

  • You will own the linguistic foundation of Innodata's segmentation and transcription work across languages, domains, and use cases. Concretely, you will: 
  • Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, timestamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions. 
  • Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale. 
  • Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale. 
  • Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can't do rather than on what they already can. 
  • Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions. 
  • Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models. 
  • Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep quality consistent as the work scales. 
  • Represent Innodata's transcription and segmentation approach to the customers and frontier labs we partner with, and contribute to the methodology and best-practice documentation that make our work legible to their teams. 

You’ll Thrive in This Role If You Have:

  • Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality — enough that you have authored standards, not only followed them. This is a principal-level role, and we weight practical depth heavily. 
  • A Bachelor's degree in linguistics, phonetics, or computational linguistics, or a closely related field, is required — with a strong foundation in phonetics, phonology, and sociolinguistics so that IPA, prosody, disfluency, dialect, and register are native concepts. An advanced degree is preferred. 
  • A big-picture grasp of how transcription and segmentation choices flow downstream into modeling — how different methods change what speech and content-understanding models learn, and therefore which method fits which modeling objective. You can explain to a model builder why a transcription decision matters. 
  • Fluency in phonetic transcription and IPA, plus hands-on experience with acoustic and phonetic analysis of speech — spectrograms, formants, pitch and prosody, and segment boundaries — applied to real, messy speech data at scale (for example in Praat). 
  • Deep experience with audio segmentation and its conventions — utterance and turn boundaries, timestamping, and speaker and diarization labeling — across real-world audio. 
  • Hands-on fluency with the modern speech stack: Whisper and the commercial ASR engines your partners benchmark against (such as AssemblyAI, Deepgram, Rev, and Speechmatics), forced alignment (for example the Montreal Forced Aligner), and annotation tools such as ELAN. 
  • Comfort scripting for speech-data work — Python for batch processing, QA, and metrics such as inter-annotator agreement and WER, plus regular expressions and Praat scripting — enough to work fluently with data and pipelines without needing an engineer for every task. 
  • Practical data-management skills across the delivery lifecycle — pre-processing, acceptance-stage quality checks, post-processing, pre-delivery validation, and report creation and packaging — so deliverables leave the door correct, consistent, and well documented. 
  • Multilingual capability and hands-on experience with accented, dialectal, and code-switched speech; low-resource languages a strong plus. 
  • A point of view on how human-in-the-loop workflows should evolve as models improve — where humans stay in the loop, where they move up to adjudication and standard-setting, and how to measure the difference. 
  • Strong written and verbal communication, comfortable working directly with research scientists and interfacing with the customers and frontier labs we partner with. 
  • Bonus: responsible-AI considerations for speech, such as bias across accents and dialects and privacy and consent in voice data. 

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Computational Linguist Related jobs

Other jobs at Innodata Inc.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.