Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Baseten

Post-Training Applied Researcher

Baseten

. Design and run post-training pipelines using SFT, GRPO, DPO, RLVR, reward function engineering, and synthetic data generation .

Posted 10/7/2026full-timeSan Francisco • California • United StatesMid-LevelSenior💰 $200,000 - $275,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in training Large Language Models (LLMs) using reinforcement learning techniques, with a strong focus on reward engineering and multi-turn agent environments. Proficient in designing and analyzing training pipelines and experiments across various domains, including healthcare and legal.

Highest-signal resume keywords
Training LLMs With Reinforcement LearningReward EngineeringMulti-Turn Agent EnvironmentsProduction ML SystemsPublications At NeurIPS, ICML, Or ICLR

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SFTGRPODPORLVRReward Function EngineeringSynthetic Data GenerationDataset ConstructionTrainingEvaluationDeployment
Tools & Technologies
RL Training FrameworksBaseten's Open-Source Training Libraries
Industry Keywords
HealthcareCode GenerationLegal DomainsReward LoopsDistribution ShiftEnd-To-End Training ExperimentsReward HackingImportance Sampling DriftAdvantage Estimation Instabilities

About the role

Key responsibilities & impact
  • Design and run post-training pipelines using SFT, GRPO, DPO, RLVR, reward function engineering, and synthetic data generation
  • Build task-specific training environments and evaluations for healthcare, code generation, and legal domains
  • Develop multi-turn tool-use, sandboxed-execution, and agentic workflows
  • Work directly with customers to translate production data into training signal
  • Design reward loops from real usage patterns and handle distribution shift
  • Run and analyze end-to-end training experiments
  • Diagnose reward hacking, importance sampling drift, and advantage estimation instabilities
  • Publish findings at top venues
  • Contribute to Baseten's open-source training libraries

Requirements

What you’ll need
  • Hands-on experience training LLMs with reinforcement learning
  • Understanding of GRPO or PPO, including group advantage computation, clipped objectives, and KL penalty design
  • Strong intuition for reward engineering
  • Experience building multi-turn agent environments with tool use
  • Comfort across dataset construction, training, evaluation, and deployment
  • Experience with production ML systems
  • Preferred: experience with RL training frameworks
  • Preferred: publications at NeurIPS, ICML, or ICLR focused on RL for LLMs, reward modeling, or alignment

Benefits

Comp & perks
  • Competitive compensation, including meaningful equity
  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • (U.S. only) Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering learning and networking opportunities