Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Handshake

Technical Staff Member – Evals

Handshake

. Design, build, and publish coding benchmarks measuring progress in frontier coding agents .

Posted 9/29/2026full-timeUnited StatesLead💰 $200,000 - $350,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates strong software engineering skills with a focus on Python and agentic software development, including experience in building coding benchmarks and evaluation methodologies. Capable of collaborating effectively in fast-paced environments while maintaining high standards of code quality and technical communication.

Highest-signal resume keywords
Strong Python SkillsExperience Designing Coding BenchmarksExpertise in Reinforcement LearningCollaborative CommunicationExperience with Automated Graders

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringCoding Benchmark DevelopmentAutomated GradersTest HarnessesProgrammatic VerifiersReinforcement LearningData WorkflowsEvaluation MethodologyHypothesis FormationMetric Selection
Soft Skills
Collaborative CommunicationLow-Ego InteractionExperimental JudgmentOwnership in Ambiguous Environments
Tools & Technologies
ML ToolingEvaluation InfrastructureOpen-Source ToolsGitHub
Industry Keywords
Agentic Software DevelopmentCoding AgentsAI for CodingCode GenerationReinforcement LearningTechnical WritingResearch Prototypes

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Design, build, and publish coding benchmarks measuring progress in frontier coding agents
  • Create realistic, difficult software-engineering tasks, repositories, environments, and test harnesses
  • Develop reliable verifiers, graders, reward signals, and evaluation methodology for agentic software development
  • Research representative, difficult, robust, and shortcut-resistant coding-agent evaluations
  • Analyze coding-agent behavior and trajectories to understand failures, useful feedback, and important capabilities
  • Partner with AI researchers, software engineers, and expert contributors on high-signal tasks, data, and evaluation methods
  • Run rapid iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and develop the next benchmark or system
  • Productize repeatable patterns into reusable software, benchmarks, datasets, and platforms
  • Raise the technical bar through design judgment, communication, code quality, and mentorship
  • Contribute open benchmarks, open-source tools, research, and technical writing
  • Help define how frontier coding agents are measured, understood, and improved

Requirements

What you’ll need
  • Deep enthusiasm for agentic software development, with evidence of actively building, experimenting with, or seriously thinking about coding agents
  • Strong software engineering skills and ability to write clean, reliable, maintainable code
  • Strong Python skills
  • Comfort with modern ML tooling, evaluation infrastructure, and data workflows
  • Sound experimental judgment, including hypothesis formation, metric selection, failure diagnosis, and distinguishing genuine capability improvement from evaluation artifacts
  • Experience designing systems and making tradeoffs around validity, quality, scale, reliability, and reuse
  • Comfort operating in an ambiguous, fast-moving environment with substantial ownership
  • Collaborative, low-ego communication and ability to work with researchers, engineers, domain experts, and customers
  • At least one of: experience working on a widely used coding-AI benchmark or evaluation suite; published research on AI for coding or code generation at a leading venue such as NeurIPS, ICML, ICLR, or COLM; software engineering experience at a top technology company or a strong public GitHub profile paired with deep knowledge of agentic AI and passion for building software with agents
  • Experience building or maintaining coding benchmarks, coding-agent environments, repository-level evaluation suites, or open-source developer tools is especially compelling
  • Experience developing automated graders, test harnesses, programmatic verifiers, reward models, or reinforcement-learning environments for software tasks is especially compelling
  • Research experience in code generation, program synthesis, AI agents, reinforcement learning, post-training, or software-engineering productivity is especially compelling
  • Experience with reinforcement learning, RLHF, preference optimization, supervised fine-tuning, or reward modeling is especially compelling
  • Strong public contributions through GitHub, papers, benchmarks, technical writing, or developer communities are especially compelling
  • Experience turning research prototypes or repeated customer work into robust, reusable products or platforms is especially compelling

Benefits

Comp & perks
  • Equity in a fast-growing company
  • 401(k) match
  • Competitive compensation
  • Financial coaching
  • Paid parental leave
  • Fertility benefits
  • Parental coaching
  • Medical, dental, and vision insurance
  • Mental health support
  • $500 wellness stipend
  • $2,000 learning stipend
  • Ongoing development
  • Commuting support
  • Free lunch
  • Gym in the San Francisco office
  • Flexible PTO
  • 15 holidays + 2 flex days
  • Team outings
  • Referral bonuses