Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Deepgram

Research Staff – Voice AI Foundations

Deepgram

. Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction .

Posted 10/11/2026full-timeRemote • AustraliaLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing neural audio codecs and generative models, with a strong foundation in statistical learning theory and experience in optimizing models for real-world deployment. Proven ability to design experiments and curate large datasets while advancing research in voice AI and multimodal learning.

Highest-signal resume keywords
Neural Audio Codec DevelopmentGenerative Model ArchitectureStatistical Learning TheoryMultimodal LearningReal-World Model Optimization

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Neural Audio CodecsGenerative ModelsStatistical Learning TheoryMultimodal Speech-to-Speech SystemsModel Architecture DesignTraining Scheme DevelopmentInference Algorithm OptimizationDataset CurationMathematical FormulationControlled Experiment Design
Soft Skills
Critical ThinkingProblem SolvingCollaboration
Tools & Technologies
AI ToolsBare-Metal Hardware
Industry Keywords
Voice AISpeech ModelingLanguage ModelingLatent RepresentationsEmbedding Systems

About the role

Key responsibilities & impact
  • Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
  • Pioneer steerable generative models for diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
  • Develop embedding systems that factorize latent representations into speaker, content, style, environment, and channel dimensions
  • Use latent recombination to generate synthetic audio data at previously impossible scales
  • Train multimodal speech-to-speech systems for robust understanding and empathic human-like responses
  • Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
  • Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
  • Conduct rigorous controlled experiments, ablations, evaluations, stress tests, and edge-case analyses
  • Contribute to foundational research advancing voice AI, speech, and language modeling

Requirements

What you’ll need
  • Strong mathematical foundation in statistical learning theory, particularly self-supervised and multimodal learning
  • Deep expertise in foundation model architectures and scaling training across multiple modalities
  • Ability to derive novel mathematical formulations and implement them efficiently
  • Demonstrated ability to build and curate massive datasets while maintaining quality and diversity
  • Track record of designing controlled experiments to validate architectural innovations and theoretical insights
  • Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
  • History of open-source contributions or research publications advancing speech/language AI
  • Comfort actively using and experimenting with advanced AI tools
  • Ability to identify critical experiments that validate or disprove ideas quickly
  • Vision to scale successful proofs-of-concept 100x
  • Experience or knowledge relevant to neural audio codecs, generative models, embedding systems, latent representations, multimodal speech-to-speech systems, model architectures, training schemes, inference algorithms, and hardware-efficient deployment

Benefits

Comp & perks
  • Remote work arrangement
  • AI-first work environment with encouragement to use and experiment with advanced AI tools
  • Opportunity to work on transformative voice AI research at scale
  • Open-source contribution and research publication opportunities
  • AI Notetaker interview recording can be declined without affecting candidacy