FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Researcher, Evaluations and Benchmarks
ActiveFence. Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI safety and security evaluations, with a strong background in building taxonomies and evaluation harnesses. Proven ability to manage research timelines, collaborate with AI labs, and communicate complex ideas effectively to diverse audiences.
Highest-signal resume keywords
PhD Or Masters In Computer Science3+ Years Building And Running Safety Evaluations5+ Relevant Research Publications In AI SafetyStrong Engineering Skills Including Evaluation HarnessesStrong Verbal And Written Communication
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Safety EvaluationsTaxonomy DevelopmentEvaluation HarnessesDistributed InferenceVLLMResearch PublicationSFTDPOGRPOAgentic Evaluation
Soft Skills
Curiosity About AI HarmsAbility To Direct FreelancersPresentation Skills
Industry Keywords
AI SafetyAI SecurityLanguage ModelsResearch CollaborationBenchmarkingQuality StandardsClient CommunicationConferences
About the role
Key responsibilities & impact- Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk
- Collaborate with leading AI labs and universities on benchmarks and papers
- Own the taxonomy, evaluation harness, quality bar, and release for each benchmark
- Review evals personally and ensure verifiers, rubrics, distributions, and taxonomies meet quality standards
- Hold the benchmark plan and calendar and keep researchers on timeline
- Direct two or three ad-hoc SME freelancers
- Set the quarterly release roadmap with the CTO, pod, and research leads based on research, client requests, and current events
- Read relevant research and maintain relationships with AI labs
- Speak with lab contacts weekly and attend conferences
Requirements
What you’ll need- PhD or Masters in computer science, machine learning or a related field, or equivalent depth from industry research
- 3+ years building and running safety or security evaluations for language models in production, at an AI lab, a model provider, or a safety and security research organisation
- 5+ relevant research publications in AI safety and security, including lead author on at least 2
- Strong engineering skills, including evaluation harnesses, distributed inference, vLLM, reading a codebase and fixing it
- Ability to build a taxonomy, not only score against one
- Ability to direct a researcher and two freelancers without formally managing them
- Strong English, written and spoken
- Curiosity about AI harms and ability to learn a new subject every three weeks
- Ideally: post-training experience with SFT, DPO, and GRPO
- Ideally: reward design for subjective and safety-relevant targets
- Ideally: agentic evaluation experience with tool use, orchestration, permissions, and prompt injection
- Ideally: publications at top conferences
- Willingness to present work on client calls
- Strong verbal and written communication and ability to present to large and/or senior audiences
- Travel to conferences at least 3 times a year
Benefits
Comp & perks- Travel to a couple of conferences a year
- Travel to conferences at least 3 times a year
- Opportunity to collaborate with leading AI labs and universities
- Opportunity to work with approximately 150 researchers on AI harms
- Role in the CTO office with access to research teams and freelancers