Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
hud (YC W25)

Research Manager

hud (YC W25)

. Lead research that makes HUD’s agent training data and evals more useful for improving frontier models .

Posted 9/26/2026full-timeSan Francisco • California • United StatesMid-LevelSeniorWebsite

About the role

Key responsibilities & impact
  • Lead research that makes HUD’s agent training data and evals more useful for improving frontier models
  • Lead research engineers through ambiguous projects and turn findings into methods usable at scale
  • Set research direction for data quality, including measuring whether tasks, trajectories, rewards, and evals are reliable and useful for training agents
  • Lead research engineers from problem definition through experiments, implementation, and clear conclusions
  • Coach research engineers to strengthen technical judgment and execution
  • Design and review experiments connecting model behavior and failure modes to data, environment, and reward design
  • Develop methods for validating and improving training data at scale, including trajectory audits, grader checks, and feedback loops
  • Partner with research engineers, domain experts, and data vendors to turn research insights into better workflows, tools, and quality standards
  • Communicate findings and tradeoffs clearly so the team can prioritize work and apply learning across research areas

Requirements

What you’ll need
  • Experience leading technical research projects to completion, from an open question to evidence, a decision, and a working result
  • Experience directly managing and mentoring researchers or research engineers while remaining engaged in technical work
  • Strong understanding of machine learning and reinforcement learning, including how training objectives, data, and feedback shape model behavior
  • Experience with agent training data, evals, benchmarks, synthetic data, or model evaluation infrastructure
  • Sound experimental judgment, including distinguishing useful training signals from tasks or metrics that only look convincing
  • Strong written communication and ability to explain methods and findings to researchers, engineers, and external partners
  • Strong candidates may also have built scalable data quality systems or validation pipelines for model training
  • Strong candidates may also have experience diagnosing reward hacking, grader errors, or other subtle agent failure modes
  • Strong candidates may also have experience translating research findings into tools and processes used by others
  • Strong candidates may also have early-stage startup experience and strong communication skills for collaboration across teams and time zones
  • Motivated candidates are encouraged to apply even if they do not meet all criteria
  • Ability to work hours that 70-80% overlap with either San Francisco or Singapore time zones
  • At least one of a LinkedIn profile or resume is required for application submission
  • Evidence of quantitative or analytical ability is required

Benefits

Comp & perks
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office (in-office employees)
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Equinox membership
  • 401k
  • Commuter benefits (US employees)
  • Unlimited access to tokens for ChatGPT, Claude Code, Cursor, etc.
  • Support for relocation and visas for strong full-time candidates to the US or Singapore