FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in reinforcement learning (RL) post-training for video models, with a strong focus on preference optimization, reward model development, and multi-node distributed training. Proficient in using PyTorch and familiar with video-specific challenges to enhance generative modeling outcomes.
Highest-signal resume keywords
Reinforcement Learning Post-TrainingPreference OptimizationPyTorch ProficiencyMulti-Node Distributed TrainingReward Model Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Generative ModelingDiffusion ModelsFlow-Matching ModelsModel Improvement EvidenceDistillation ExpertiseAdversarial ApproachesEvaluation MetricsData Collection WorkflowsHuman Preference StudiesVideo-Specific Failure Modes
Soft Skills
CollaborationCommunication
Industry Keywords
Video GenerationTemporal DriftMotion RealismPhysics RealismVLM-as-Judge
Tech Stack
Tools & technologiesNode.jsPyTorch
About the role
Key responsibilities & impact- Own RL-based post-training for video diffusion/flow-matching models
- Run RL post-training, including preference optimization and online RL against learned rewards, at multi-node scale
- Build video reward models by defining evaluation dimensions, configuring preference data collection workflows, training and validating learned judges, and safeguarding against reward hacking
- Own post-training evaluation using human preference studies and correlation with automated metrics
- Distill RL-tuned models to efficient few-step samplers while preserving alignment gains
- Collaborate with engineering and product teams
- Shape real-time creative and agentic video platforms
Requirements
What you’ll need- 2+ years hands-on research experience in post-training or generative modeling
- RL or preference-optimization experience on generative models with evidence of model improvement
- Strong grounding in diffusion or flow-matching models
- Proficiency with PyTorch
- Experience with multi-node distributed training
- Experience developing reward models for visual generation, including VLM-as-judge or large-scale preference data collection (preferred)
- Distillation expertise in distribution matching, consistency, or adversarial approaches, ideally for video models (preferred)
- Familiarity with video-specific failure modes such as temporal drift, motion, and physics realism (preferred)
- Ability to regularly work from the Palo Alto HQ
- Authorization to work in the United States or ability to address US employment sponsorship requirements
Benefits
Comp & perks- Substantial equity in a high-growth startup
- Full health benefits
- 401k matching
- Significant growth opportunities
- Flexible on-site/remote hybrid work arrangement
