FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer, Synthetic Data
hud (YC W25). Build HUD’s synthetic data pipeline for RL training data and evals for frontier AI agents .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Proficient in building synthetic data pipelines and analyzing model performance, with a strong understanding of synthetic data quality and its limitations. Capable of designing diverse and realistic synthetic tasks while collaborating effectively in remote environments.
Highest-signal resume keywords
Proficiency In PythonExperience With Synthetic Data Research MethodsBuilding Synthetic Data Pipelines End-To-EndStrong Communication SkillsDetail-Oriented
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonSynthetic Data Research MethodsData Pipeline DevelopmentTask DesignPerformance AnalysisMetrics Development
Soft Skills
Detail-OrientedStrong Communication SkillsIndependent Work
Tools & Technologies
DockerLinux
Industry Keywords
Synthetic DataReinforcement LearningTask Learning OutcomesEvaluation BenchmarksEarly-Stage Startup Experience
Tech Stack
Tools & technologiesDockerLinuxPython
About the role
Key responsibilities & impact- Build HUD’s synthetic data pipeline for RL training data and evals for frontier AI agents
- Work with subject-matter experts to create synthetic tasks across professional and technical domains
- Design methods that generate diverse, realistic, and learnable synthetic tasks
- Build systems and tooling to mutate, validate, and improve synthetic tasks
- Analyze model and agent performance to understand task learning outcomes and failure modes
- Develop metrics for synthetic task diversity, realism, and learnability
- Complete a process involving two technical interviews and a 2–3 day work trial
Requirements
What you’ll need- Proficiency in Python, Docker, and Linux environments
- Experience with synthetic data research methods; applicants should elaborate in their application
- Strong understanding of what “good synthetic data” means and its limitations
- Experience building synthetic data pipelines end-to-end without a fully prescribed roadmap
- Experience working on environments, evals, and benchmarks
- Detail-oriented and able to spot subtle inconsistencies or edge cases in synthetic data
- Able to reason from first principles about task design, scoring, and failure modes
- Early-stage startup experience and ability to work independently in fast-paced environments
- Strong communication skills for remote collaboration across time zones
- Motivated candidates are encouraged to apply even if they do not meet all criteria
Benefits
Comp & perks- 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
- Lunch and dinner when you’re in the office (in-office employees)
- Company-wide holiday break from Christmas Eve to New Year’s Day on top of PTO and paid holidays
- Equinox membership (US employees)
- 401k (US employees)
- Commuter benefits (US employees)
- Unlimited access to tokens for ChatGPT, Claude Code, Cursor, etc.
- Support for relocation and visas for strong full-time candidates to the US or Singapore