Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Pika

Research Scientist – Video Generation

Pika

. Own RL-based post-training for video diffusion/flow-matching models .

Posted 9/24/2026full-timePalo Alto • California • United StatesJuniorMid-Level💰 $185,000 - $400,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in reinforcement learning (RL) post-training for video models, with a strong focus on preference optimization, reward model development, and multi-node distributed training. Proficient in using PyTorch and familiar with video-specific challenges to enhance generative modeling outcomes.

Highest-signal resume keywords
Reinforcement Learning Post-TrainingPreference OptimizationPyTorch ProficiencyMulti-Node Distributed TrainingReward Model Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Generative ModelingDiffusion ModelsFlow-Matching ModelsModel Improvement EvidenceDistillation ExpertiseAdversarial ApproachesEvaluation MetricsData Collection WorkflowsHuman Preference StudiesVideo-Specific Failure Modes
Soft Skills
CollaborationCommunication
Industry Keywords
Video GenerationTemporal DriftMotion RealismPhysics RealismVLM-as-Judge

Tech Stack

Tools & technologies
Node.jsPyTorch

About the role

Key responsibilities & impact
  • Own RL-based post-training for video diffusion/flow-matching models
  • Run RL post-training, including preference optimization and online RL against learned rewards, at multi-node scale
  • Build video reward models by defining evaluation dimensions, configuring preference data collection workflows, training and validating learned judges, and safeguarding against reward hacking
  • Own post-training evaluation using human preference studies and correlation with automated metrics
  • Distill RL-tuned models to efficient few-step samplers while preserving alignment gains
  • Collaborate with engineering and product teams
  • Shape real-time creative and agentic video platforms

Requirements

What you’ll need
  • 2+ years hands-on research experience in post-training or generative modeling
  • RL or preference-optimization experience on generative models with evidence of model improvement
  • Strong grounding in diffusion or flow-matching models
  • Proficiency with PyTorch
  • Experience with multi-node distributed training
  • Experience developing reward models for visual generation, including VLM-as-judge or large-scale preference data collection (preferred)
  • Distillation expertise in distribution matching, consistency, or adversarial approaches, ideally for video models (preferred)
  • Familiarity with video-specific failure modes such as temporal drift, motion, and physics realism (preferred)
  • Ability to regularly work from the Palo Alto HQ
  • Authorization to work in the United States or ability to address US employment sponsorship requirements

Benefits

Comp & perks
  • Substantial equity in a high-growth startup
  • Full health benefits
  • 401k matching
  • Significant growth opportunities
  • Flexible on-site/remote hybrid work arrangement