FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in generative AI and multimodal models, with a strong focus on dyadic interaction modeling and video diffusion techniques. Proficient in building evaluation frameworks and conducting rigorous experiments to enhance interactive agent performance.
Highest-signal resume keywords
Strong ML BackgroundHands-On Experience With Diffusion ModelsProficiency In PyTorchPublications At Top-Tier VenuesExperience With Real-Time Generation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Diffusion ModelsMultimodal ModelsDyadic Interaction ModelingVideo GenerationLarge-Scale TrainingAutoregressive Video DiffusionLow-Latency Inference TechniquesAudio-Driven Motion ModelingConversational ModelingResearch Implementation
Soft Skills
Clear CommunicationMentoring Experience
Tools & Technologies
PyTorchModern ML Tooling
Industry Keywords
Generative AIInteractive VideoEvaluation FrameworksData Needs DefinitionTechnical Decision Making
Tech Stack
Tools & technologiesPyTorch
About the role
Key responsibilities & impact- Join a team of 40+ researchers and engineers in R&D working on generative AI, avatar-centric interactive video diffusion models
- Contribute to research direction for dyadic interaction modeling and own well-scoped research problems end to end
- Advance the perceptual layer of interactive agents by understanding user audio and video and generating contextually appropriate reactions
- Post-train multimodal models to generate natural dyadic interactions from user audio and video inputs
- Adapt diffusion models to conditioning signals such as conversational state, turn-taking, and listener cues
- Build evaluation frameworks and test suites for tracking interaction quality
- Partner with the data team to define data needs and shape high-quality datasets
- Run rigorous experiments and share findings that inform technical decisions
Requirements
What you’ll need- Strong ML background and hands-on experience with diffusion models, ideally for video or avatar generation
- Publications at top-tier venues such as CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or SIGGRAPH in world models, dyadic interaction, or video diffusion, or equivalent demonstrated impact
- Experience taking research ideas through to working implementations
- Proficiency in PyTorch and modern ML tooling for large-scale training
- Clear communication of hypotheses, experiments, and results
- Experience with real-time or streaming generation, including autoregressive video diffusion (nice to have)
- Distillation or other techniques for low-latency inference (nice to have)
- Audio-driven facial, gesture, or full-body motion modeling (nice to have)
- Conversational modeling, such as turn-taking, backchanneling, or listener-response generation (nice to have)
- Experience mentoring students or junior researchers (nice to have)
- Ability to work legally in the country of employment without visa sponsorship, as indicated in the application form
Benefits
Comp & perks- Build production-scale video foundation models in a fast-growing Generative AI company
- Work on human-centric video generation with real-world impact
- Tackle hard problems in scaling, stability, and controllability
- Influence the direction of next-generation synthetic human technology
- Highly technical, high-ownership environment where work ships
