Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Zillow

Senior Applied Scientist

Zillow

. Own LLM post-training pipelines, including SFT, DPO, and reinforcement fine-tuning (RFT/GRPO), running training end-to-end on GPU infrastructure .

Posted 9/17/2026full-timeRemote • United StatesSenior💰 $152,900 - $257,100 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in LLM post-training pipelines, including SFT, DPO, and reinforcement fine-tuning, with a strong foundation in generative AI and reward model development. Capable of providing technical leadership and mentorship while driving complex projects in collaboration with cross-functional teams.

Highest-signal resume keywords
LLM Post-Training PipelinesReinforcement Learning Fine-TuningGenerative AI KnowledgePython ProgrammingML Frameworks (PyTorch, TensorFlow)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SFTDPORFTGRPOReward Model DevelopmentEnd-to-End TrainingMulti-Turn Trace AnalysisLLM-as-a-JudgePreference LearningGPU Infrastructure
Soft Skills
Technical LeadershipMentorshipProblem SolvingCollaboration
Tools & Technologies
DatabricksFireworks
Certifications & Qualifications
Master's Degree in Computer SciencePhD in Computer Science or Machine Learning
Industry Keywords
Generative AIFoundation ModelsTransformersReinforcement LearningPreference Learning

Tech Stack

Tools & technologies
PythonPyTorchTensorflow

About the role

Key responsibilities & impact
  • Own LLM post-training pipelines, including SFT, DPO, and reinforcement fine-tuning (RFT/GRPO), running training end-to-end on GPU infrastructure
  • Build and train fast, multi-category reward models that emit scalar signals from contrastive preference pairs
  • Design on-policy and online assessment by scoring live agent trajectories using LLM-as-a-Judge frameworks
  • Translate offline evaluation rubrics into generalizable reward functions for arbitrary production traces
  • Drive complex, ambiguous work end-to-end in collaboration with product, engineering, science, platform, and data teams
  • Provide technical leadership and mentorship
  • Translate state-of-the-art AI research into production impact

Requirements

What you’ll need
  • Master's degree or higher in Computer Science or a related field
  • Hands-on experience with LLM post-training and RL fine-tuning: SFT, DPO, RFT/GRPO, running end-to-end training runs on GPU infrastructure
  • Exposure to reward model development; research-level experience is sufficient
  • Strong knowledge of generative AI, including foundation models, transformers, reinforcement learning, and preference learning
  • Ability to independently scope and solve ambiguous problems end-to-end while providing technical leadership to scientists and MLEs
  • Strong programming skills, especially Python
  • Experience with ML frameworks such as PyTorch or TensorFlow
  • PhD in Computer Science, Machine Learning, or a related field (nice-to-have)
  • Experience evaluating agentic AI: multi-turn trace analysis, LLM-as-a-Judge, and the interaction between offline evals and online monitoring (nice-to-have)
  • Published work in post-training, RLHF/RLAIF, preference learning, or reward modeling (nice-to-have)
  • Experience with GPU training platforms such as Databricks or Fireworks (nice-to-have)

Benefits

Comp & perks
  • Equity awards based on factors such as experience, performance and location
  • Remote work from a physical location of choice
  • Flexibility to work from wherever most productive
  • Accommodation support for disabilities or special needs