FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer – Benchmarks
hud (YC W25). Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates proficiency in Python, Docker, and Linux environments to build and evaluate domain-specific benchmarks. Capable of designing reliable metrics and analyses while effectively communicating findings to technical audiences.
Highest-signal resume keywords
Proficiency In PythonExperience With DockerStrong Understanding Of Benchmark DesignDetail-Oriented AnalysisStrong Communication Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonBenchmark DesignMetric DevelopmentTask EvaluationAnalytical Skills
Soft Skills
CuriosityDetail-OrientedAbility To Thrive In Unstructured EnvironmentsStrong Communication Skills
Tools & Technologies
DockerLinux
Industry Keywords
BenchmarkingModel EvaluationReal-World CorrelationTechnical DocumentationStartup Experience
Tech Stack
Tools & technologiesDockerLinuxPython
About the role
Key responsibilities & impact- Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks
- Own the design, implementation, and quality of HUD’s internal agent benchmarks
- Work with subject-matter experts to define tasks and create domain-specific benchmarks evaluating realistic workflows
- Build infrastructure to reliably run models and agents against benchmark tasks
- Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
- Validate whether benchmark performance correlates with real-world evaluations, customer needs, and lab expectations
- Write clear documentation and benchmark reports for technical audiences
- Participate in a 2–3 day work trial during the hiring process
Requirements
What you’ll need- Proficiency in Python, Docker, and Linux environments
- Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc.; link in application
- Strong understanding of what makes a benchmark realistic, reliable, and useful
- Experience working on environments and evals
- Curiosity and ability to understand workflows in various domains
- Detail-oriented and able to spot subtle inconsistencies or edge cases in tasks
- Ability to reason from first principles about task design, scoring, and failure modes
- Ability to thrive in unstructured problem spaces
- Early-stage startup experience and ability to work independently in fast-paced environments
- Strong communication skills for remote collaboration across time zones
- Ability to work hours that 70–80% overlap with either San Francisco or Singapore time zones for remote work
- Technical aptitude and learning potential; years of experience are not required
Benefits
Comp & perks- 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
- Lunch and dinner when you’re in the office (in-office employees)
- Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
- PTO and paid holidays
- Equinox membership (US employees)
- 401k (US employees)
- Commuter benefits (US employees)
- Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc.
- Support for relocation and visas for strong full-time candidates to the US or Singapore
- Competitive compensation