FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in LLM post-training pipelines, including SFT, DPO, and reinforcement fine-tuning, with a strong foundation in generative AI and reward model development. Capable of providing technical leadership and mentorship while driving complex projects in collaboration with cross-functional teams.
Highest-signal resume keywords
LLM Post-Training PipelinesReinforcement Learning Fine-TuningGenerative AI KnowledgePython ProgrammingML Frameworks (PyTorch, TensorFlow)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SFTDPORFTGRPOReward Model DevelopmentEnd-to-End TrainingMulti-Turn Trace AnalysisLLM-as-a-JudgePreference LearningGPU Infrastructure
Soft Skills
Technical LeadershipMentorshipProblem SolvingCollaboration
Tools & Technologies
DatabricksFireworks
Certifications & Qualifications
Master's Degree in Computer SciencePhD in Computer Science or Machine Learning
Industry Keywords
Generative AIFoundation ModelsTransformersReinforcement LearningPreference Learning
Tech Stack
Tools & technologiesPythonPyTorchTensorflow
About the role
Key responsibilities & impact- Own LLM post-training pipelines, including SFT, DPO, and reinforcement fine-tuning (RFT/GRPO), running training end-to-end on GPU infrastructure
- Build and train fast, multi-category reward models that emit scalar signals from contrastive preference pairs
- Design on-policy and online assessment by scoring live agent trajectories using LLM-as-a-Judge frameworks
- Translate offline evaluation rubrics into generalizable reward functions for arbitrary production traces
- Drive complex, ambiguous work end-to-end in collaboration with product, engineering, science, platform, and data teams
- Provide technical leadership and mentorship
- Translate state-of-the-art AI research into production impact
Requirements
What you’ll need- Master's degree or higher in Computer Science or a related field
- Hands-on experience with LLM post-training and RL fine-tuning: SFT, DPO, RFT/GRPO, running end-to-end training runs on GPU infrastructure
- Exposure to reward model development; research-level experience is sufficient
- Strong knowledge of generative AI, including foundation models, transformers, reinforcement learning, and preference learning
- Ability to independently scope and solve ambiguous problems end-to-end while providing technical leadership to scientists and MLEs
- Strong programming skills, especially Python
- Experience with ML frameworks such as PyTorch or TensorFlow
- PhD in Computer Science, Machine Learning, or a related field (nice-to-have)
- Experience evaluating agentic AI: multi-turn trace analysis, LLM-as-a-Judge, and the interaction between offline evals and online monitoring (nice-to-have)
- Published work in post-training, RLHF/RLAIF, preference learning, or reward modeling (nice-to-have)
- Experience with GPU training platforms such as Databricks or Fireworks (nice-to-have)
Benefits
Comp & perks- Equity awards based on factors such as experience, performance and location
- Remote work from a physical location of choice
- Flexibility to work from wherever most productive
- Accommodation support for disabilities or special needs
