FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in training Large Language Models (LLMs) using reinforcement learning techniques, with a strong focus on reward engineering and multi-turn agent environments. Proficient in designing and analyzing training pipelines and experiments across various domains, including healthcare and legal.
Highest-signal resume keywords
Training LLMs With Reinforcement LearningReward EngineeringMulti-Turn Agent EnvironmentsProduction ML SystemsPublications At NeurIPS, ICML, Or ICLR
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SFTGRPODPORLVRReward Function EngineeringSynthetic Data GenerationDataset ConstructionTrainingEvaluationDeployment
Tools & Technologies
RL Training FrameworksBaseten's Open-Source Training Libraries
Industry Keywords
HealthcareCode GenerationLegal DomainsReward LoopsDistribution ShiftEnd-To-End Training ExperimentsReward HackingImportance Sampling DriftAdvantage Estimation Instabilities
About the role
Key responsibilities & impact- Design and run post-training pipelines using SFT, GRPO, DPO, RLVR, reward function engineering, and synthetic data generation
- Build task-specific training environments and evaluations for healthcare, code generation, and legal domains
- Develop multi-turn tool-use, sandboxed-execution, and agentic workflows
- Work directly with customers to translate production data into training signal
- Design reward loops from real usage patterns and handle distribution shift
- Run and analyze end-to-end training experiments
- Diagnose reward hacking, importance sampling drift, and advantage estimation instabilities
- Publish findings at top venues
- Contribute to Baseten's open-source training libraries
Requirements
What you’ll need- Hands-on experience training LLMs with reinforcement learning
- Understanding of GRPO or PPO, including group advantage computation, clipped objectives, and KL penalty design
- Strong intuition for reward engineering
- Experience building multi-turn agent environments with tool use
- Comfort across dataset construction, training, evaluation, and deployment
- Experience with production ML systems
- Preferred: experience with RL training frameworks
- Preferred: publications at NeurIPS, ICML, or ICLR focused on RL for LLMs, reward modeling, or alignment
Benefits
Comp & perks- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- (U.S. only) Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering learning and networking opportunities
