FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing neural audio codecs and generative models, with a strong foundation in statistical learning theory and experience in optimizing models for real-world deployment. Proven ability to design experiments and curate large datasets while advancing research in voice AI and multimodal learning.
Highest-signal resume keywords
Neural Audio Codec DevelopmentGenerative Model ArchitectureStatistical Learning TheoryMultimodal LearningReal-World Model Optimization
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Neural Audio CodecsGenerative ModelsStatistical Learning TheoryMultimodal Speech-to-Speech SystemsModel Architecture DesignTraining Scheme DevelopmentInference Algorithm OptimizationDataset CurationMathematical FormulationControlled Experiment Design
Soft Skills
Critical ThinkingProblem SolvingCollaboration
Tools & Technologies
AI ToolsBare-Metal Hardware
Industry Keywords
Voice AISpeech ModelingLanguage ModelingLatent RepresentationsEmbedding Systems
About the role
Key responsibilities & impact- Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
- Pioneer steerable generative models for diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
- Develop embedding systems that factorize latent representations into speaker, content, style, environment, and channel dimensions
- Use latent recombination to generate synthetic audio data at previously impossible scales
- Train multimodal speech-to-speech systems for robust understanding and empathic human-like responses
- Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
- Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
- Conduct rigorous controlled experiments, ablations, evaluations, stress tests, and edge-case analyses
- Contribute to foundational research advancing voice AI, speech, and language modeling
Requirements
What you’ll need- Strong mathematical foundation in statistical learning theory, particularly self-supervised and multimodal learning
- Deep expertise in foundation model architectures and scaling training across multiple modalities
- Ability to derive novel mathematical formulations and implement them efficiently
- Demonstrated ability to build and curate massive datasets while maintaining quality and diversity
- Track record of designing controlled experiments to validate architectural innovations and theoretical insights
- Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
- History of open-source contributions or research publications advancing speech/language AI
- Comfort actively using and experimenting with advanced AI tools
- Ability to identify critical experiments that validate or disprove ideas quickly
- Vision to scale successful proofs-of-concept 100x
- Experience or knowledge relevant to neural audio codecs, generative models, embedding systems, latent representations, multimodal speech-to-speech systems, model architectures, training schemes, inference algorithms, and hardware-efficient deployment
Benefits
Comp & perks- Remote work arrangement
- AI-first work environment with encouragement to use and experiment with advanced AI tools
- Opportunity to work on transformative voice AI research at scale
- Open-source contribution and research publication opportunities
- AI Notetaker interview recording can be declined without affecting candidacy
