Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
InnoData

Research Scientist, Speech & Audio

InnoData

. Define how Innodata designs, structures, and evaluates audio data for speech and audio models .

Posted 9/24/2026full-timeRemote • United StatesMid-LevelSenior💰 $160,000 - $185,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and evaluating audio data for speech models, with a strong foundation in machine learning methodologies and hands-on experience in building and fine-tuning various speech and audio models. Proficient in collaborating with cross-functional teams to operationalize specifications and publish research findings.

Highest-signal resume keywords
Speech Or Audio ML ExperiencePyTorch FundamentalsFluency In ESPnet, NeMo, SpeechBrain, KaldiFirst-Author Publications In Interspeech, ICASSP, ASRUExperience With Multilingual And Accented Speech

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Speech RecognitionText-To-SpeechSpeech-To-Speech ModelsDiarizationEvaluation MethodologiesData SpecificationsDataset Coverage ReasoningNoise And Room-Response SimulationTranscription SpecificationsAdversarial Evaluations
Soft Skills
Clear CommunicationCollaborationRigorous DocumentationCustomer EngagementProblem-Solving
Tools & Technologies
HuggingFaceForced AlignmentWER/CER MetricsSynthetic Audio PipelinesAugmented Audio Pipelines
Industry Keywords
ASRConversational VoiceSpeech NaturalnessStreaming LatencyCode-SwitchingDiarization Error RateSemantic AccuracyEmotional RangeParalinguistic RangeLow-Resource Speech

Tech Stack

Tools & technologies
PyTorchSwitching

About the role

Key responsibilities & impact
  • Define how Innodata designs, structures, and evaluates audio data for speech and audio models
  • Translate speech and audio model requirements into data specifications covering modalities, transcription and annotation schemas, sampling, and evaluation criteria
  • Build evaluation methodologies covering semantic accuracy, noise and accent robustness, code-switching, diarization error rate, speech naturalness and intelligibility, and streaming latency
  • Structure, enrich, and sample audio across languages, accents, acoustic conditions, demographics, emotional and paralinguistic range, scripted and spontaneous speech, and speaker configurations
  • Partner with transcription and linguistics leads to develop transcription specifications and quantify their impact on model results
  • Partner with audio solutions and engineering teams to ensure collected audio meets model objectives
  • Run fine-tuning and evaluation experiments with ablations linking data choices to measurable improvements
  • Design adversarial and stumping evaluations to identify speech-system failures
  • Publish benchmarks, methodology, and research papers
  • Work with annotation teams, subject-matter experts, and synthetic- and augmented-audio pipelines to operationalize specifications
  • Partner directly with customers and frontier labs building ASR, text-to-speech, speech-to-speech, conversational voice, diarization, and audio-language models

Requirements

What you’ll need
  • Roughly 5+ years of hands-on industry experience in speech or audio ML
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required
  • Trained and evaluated speech or audio models yourself — ASR, TTS, speech-to-speech, speaker, or audio-language models — with strong PyTorch fundamentals
  • Fluency in ESPnet, NeMo, SpeechBrain, Kaldi, HuggingFace, forced alignment, WER/CER, and related speech metrics
  • Hands-on experience with multilingual, accented, dialectal, low-resource, or code-switched speech
  • Experience with synthetic or augmented audio, including TTS pipelines, noise and room-response simulation
  • Built evaluation sets and reasoned about dataset coverage across conditions
  • First-author publications or strong open-source contributions at venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS
  • Ability to work directly with customer and frontier-lab research scientists
  • Ability to explain data and modeling decisions clearly to expert and non-expert audiences
  • Rigorous, reproducible approach to experiments and documentation
  • Interest or hands-on experience in responsible-AI evaluation and red-teaming is a bonus