Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Deepgram

Research Staff, Voice AI Foundations

Deepgram

. Pioneer Latent Space Models to address fundamental data, scale, and cost challenges in robust, contextualized voice AI .

Posted 10/11/2026full-timeRemote • NetherlandsLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and optimizing neural audio codecs and generative models for voice AI, with a strong foundation in statistical learning theory and experience in managing large-scale data pipelines. Proven ability to conduct rigorous experiments and contribute to open-source projects in speech processing.

Highest-signal resume keywords
Neural Audio Codec DevelopmentStatistical Learning TheoryMultimodal LearningData Pipeline ManagementOpen-Source Contributions

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Latent Space ModelsGenerative ModelsSpeech ProcessingModel Architecture DesignTraining Scheme DevelopmentInference Algorithm OptimizationReal-Time InferenceControlled Experiment DesignMathematical FormulationAI Automation
Soft Skills
CollaborationCritical ThinkingProblem SolvingVisionary Thinking
Tools & Technologies
Bare-Metal HardwareHigh-Fidelity ReconstructionLow-Bit-Rate CompressionSpeech-To-Speech Systems
Industry Keywords
Speech AISelf-Supervised LearningMultimodal SystemsSTT/TTSEmpathic Responses

About the role

Key responsibilities & impact
  • Pioneer Latent Space Models to address fundamental data, scale, and cost challenges in robust, contextualized voice AI
  • Build next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
  • Develop steerable generative models synthesizing diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
  • Create embedding systems that factorize latent spaces into speaker, content, style, environment, and channel dimensions
  • Use latent recombination to generate synthetic audio data at previously impossible scales
  • Train multimodal speech-to-speech systems for universal understanding and empathic, human-like responses
  • Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
  • Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
  • Conduct rigorous controlled experiments, ablations, evaluations, and stress tests
  • Collaborate through research publications and open-source contributions

Requirements

What you’ll need
  • Strong mathematical foundation in statistical learning theory, particularly for self-supervised and multimodal learning
  • Deep expertise in foundation model architectures and scaling training across multiple modalities
  • Ability to derive novel mathematical formulations and implement them efficiently
  • Demonstrated ability to build and maintain massive, high-quality, diverse data pipelines
  • Experience designing controlled experiments to isolate architectural impacts and validate theoretical insights
  • Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
  • Track record of open-source contributions or research publications advancing speech/language AI
  • Ability to identify critical experiments that validate or disprove ideas quickly
  • Vision to scale successful proofs-of-concept 100x
  • Strong AI automation and augmentation mindset
  • Scholarly literature or academic publications related to speech processing, STT/TTS, or similar requested in the application
  • Legal authorization to work in the country where the role is located
  • Ability to address visa sponsorship requirements for the country where the role is located

Benefits

Comp & perks
  • Remote work arrangement
  • AI-first work environment with opportunities to use and experiment with advanced AI tools
  • Opportunity to work on transformative voice AI research
  • Opportunity to contribute to open-source projects and research publications
  • AI Notetaker interview recording is optional; opting out does not impact candidacy