Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
hud (YC W25)

Research Engineer – Benchmarks

hud (YC W25)

. Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks .

Posted 9/23/2026full-timeSan Francisco • California • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates proficiency in Python, Docker, and Linux environments to build and evaluate domain-specific benchmarks. Capable of designing reliable metrics and analyses while effectively communicating findings to technical audiences.

Highest-signal resume keywords
Proficiency In PythonExperience With DockerStrong Understanding Of Benchmark DesignDetail-Oriented AnalysisStrong Communication Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonBenchmark DesignMetric DevelopmentTask EvaluationAnalytical Skills
Soft Skills
CuriosityDetail-OrientedAbility To Thrive In Unstructured EnvironmentsStrong Communication Skills
Tools & Technologies
DockerLinux
Industry Keywords
BenchmarkingModel EvaluationReal-World CorrelationTechnical DocumentationStartup Experience

Tech Stack

Tools & technologies
DockerLinuxPython

About the role

Key responsibilities & impact
  • Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks
  • Own the design, implementation, and quality of HUD’s internal agent benchmarks
  • Work with subject-matter experts to define tasks and create domain-specific benchmarks evaluating realistic workflows
  • Build infrastructure to reliably run models and agents against benchmark tasks
  • Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
  • Validate whether benchmark performance correlates with real-world evaluations, customer needs, and lab expectations
  • Write clear documentation and benchmark reports for technical audiences
  • Participate in a 2–3 day work trial during the hiring process

Requirements

What you’ll need
  • Proficiency in Python, Docker, and Linux environments
  • Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc.; link in application
  • Strong understanding of what makes a benchmark realistic, reliable, and useful
  • Experience working on environments and evals
  • Curiosity and ability to understand workflows in various domains
  • Detail-oriented and able to spot subtle inconsistencies or edge cases in tasks
  • Ability to reason from first principles about task design, scoring, and failure modes
  • Ability to thrive in unstructured problem spaces
  • Early-stage startup experience and ability to work independently in fast-paced environments
  • Strong communication skills for remote collaboration across time zones
  • Ability to work hours that 70–80% overlap with either San Francisco or Singapore time zones for remote work
  • Technical aptitude and learning potential; years of experience are not required

Benefits

Comp & perks
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office (in-office employees)
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • PTO and paid holidays
  • Equinox membership (US employees)
  • 401k (US employees)
  • Commuter benefits (US employees)
  • Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc.
  • Support for relocation and visas for strong full-time candidates to the US or Singapore
  • Competitive compensation