FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing neural audio codecs and multimodal speech systems, with a strong foundation in statistical learning theory and experience in optimizing models for real-world deployment. Proven ability to design controlled experiments and contribute to open-source research in speech and language AI.
Highest-signal resume keywords
Neural Audio Codec DevelopmentStatistical Learning TheoryMultimodal LearningData Pipeline DesignOpen-Source Contributions
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Neural Audio CodecsGenerative ModelsLatent Space ModelsModel Architecture DesignStatistical Learning TheoryData Pipeline DevelopmentReal-Time Inference OptimizationControlled Experiment DesignMultimodal Speech SystemsMathematical Formulation Implementation
Soft Skills
AdaptabilityCollaborationCritical ThinkingEmpathyContinuous Learning
Industry Keywords
Speech AILanguage AISelf-Supervised LearningBillion-Hour Dataset TrainingAI Automation
About the role
Key responsibilities & impact- Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction across world-scale general audio corpora
- Pioneer steerable generative models synthesizing diverse human speech, including emotional expression, multi-speaker scenarios, environmental noise, and overlapping speech
- Develop embedding systems that factorize codec latent spaces into interpretable speaker, content, style, environment, and channel dimensions
- Use latent recombination to generate synthetic audio data at unprecedented scales
- Train multimodal speech-to-speech systems that understand diverse humans and produce empathic, human-like responses
- Design model architectures, training schemes, and inference algorithms adapted to bare-metal hardware for cost-efficient billion-hour dataset training and real-time inference
- Conduct foundational research in Latent Space Models addressing data, scale, and cost challenges in voice AI
- Design controlled experiments, ablations, evaluations, stress tests, and benchmarks to validate research hypotheses
- Collaborate through open-source contributions and research publications advancing speech and language AI
Requirements
What you’ll need- Strong mathematical foundation in statistical learning theory, particularly areas relevant to self-supervised and multimodal learning
- Deep expertise in foundation model architectures and scaling training across multiple modalities
- Proven ability to bridge theory and practice by deriving novel mathematical formulations and implementing them efficiently
- Demonstrated ability to build data pipelines that process and curate massive datasets while maintaining quality and diversity
- Track record of designing controlled experiments that isolate architectural innovations and validate theoretical insights
- Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
- History of open-source contributions or research publications advancing speech/language AI
- Ability to identify critical experiments that validate or disprove ideas quickly
- Vision to scale successful proofs-of-concept 100x
- Comfort using AI to automate and amplify personal impact
- Ability to adapt quickly, experiment, learn constantly, and work in a rapidly changing AI environment
Benefits
Comp & perks- Remote work arrangement
- AI-first work environment with active use and experimentation of advanced AI tools
- Opportunity to pioneer foundational voice AI research with transformative impact
- Opportunity to contribute to open-source projects and research publications
- AI Notetaker interview recording is optional; opting out does not impact candidacy
