FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Member of Technical Staff – Post-Training
Handshake. Design post-training systems and methodologies for frontier models, including supervised fine-tuning, reinforcement learning, preference optimization, reward modeling, and related approaches .
Posted 9/23/2026full-timeRemote • United States, Canada, United KingdomLead💰 $200,000 - $350,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing post-training systems and methodologies for machine learning models, with a strong focus on evaluation frameworks and data-processing pipelines. Proven ability to collaborate with researchers and domain experts to translate complex needs into actionable experiments and production-quality solutions.
Highest-signal resume keywords
Post-Training Systems DesignPython ProgrammingPyTorch ExperienceReinforcement LearningModel Evaluation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Supervised Fine-TuningReinforcement LearningReward ModelingData-Processing PipelinesEvaluation FrameworksQuality-Control SystemsHypothesis FormationScalable Software DevelopmentLarge-Scale ML TrainingModel Behavior Analysis
Soft Skills
Collaborative CommunicationExperimental JudgmentMentorshipAdaptabilityProblem-Solving
Tools & Technologies
ML ToolingTraining EnvironmentsEvaluation WorkflowsOpen-Source ToolsAnnotation Systems
Industry Keywords
Machine LearningAI ResearchData Quality FrameworksHuman-in-the-Loop SystemsTechnical Leadership
Tech Stack
Tools & technologiesPythonPyTorch
About the role
Key responsibilities & impact- Design post-training systems and methodologies for frontier models, including supervised fine-tuning, reinforcement learning, preference optimization, reward modeling, and related approaches
- Translate open-ended research or partner needs into hypotheses, experiments, evaluation plans, and production-quality implementations
- Build and improve evaluation frameworks, benchmarks, training environments, data-processing pipelines, and quality-control systems
- Run rapid iteration loops: prototype, evaluate, interpret results, and turn learnings into the next system or product
- Partner with AI researchers and domain experts to develop high-signal data, feedback, and evaluation methods
- Identify repeatable patterns across engagements and productize them into reusable software and platforms
- Raise the technical bar through design judgment, communication, code quality, and mentorship
- Contribute through benchmarks, open-source tools, research, and technical writing where it creates leverage
- Help define Handshake Labs' technical direction, operating culture, and reusable systems
Requirements
What you’ll need- 3+ years of demonstrated strength in post-training, fine-tuning, or model-evaluation work
- Relevant experience may include RL, SFT, LoRA/PEFT, full fine-tuning, RLHF, DPO, PPO, reward modeling, or training environments
- Strong Python skills and ability to write clean, efficient, scalable software
- Hands-on experience with modern ML tooling, particularly PyTorch and large-scale data, training, or evaluation workflows
- Sound experimental judgment, including forming hypotheses, choosing meaningful metrics, diagnosing failures, and distinguishing signal from noise
- Experience designing systems and making tradeoffs around quality, scale, reliability, and reuse
- Comfort operating in an ambiguous, fast-moving environment with substantial ownership
- Collaborative communication and ability to work with researchers, engineers, domain experts, and customers
- Especially compelling: large-scale ML training, inference, data, or evaluation systems
- Especially compelling: LLM/agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks
- Especially compelling: reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems
- Especially compelling: published research, meaningful open-source contributions, or technical leadership in ML systems or AI research
- Especially compelling: productizing research or repeated customer work into robust, reusable platforms
Benefits
Comp & perks- Equity in a fast-growing company
- 401(k) match
- Competitive compensation
- Financial coaching
- Paid parental leave
- Fertility benefits
- Parental coaching
- Medical, dental, and vision insurance
- Mental health support
- $500 wellness stipend
- $2,000 learning stipend
- Ongoing development
- Commuting support
- Free lunch
- Gym in the San Francisco office
- Flexible PTO
- 15 holidays + 2 flex days
- Team outings
- Referral bonuses