Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
ActiveFence

Researcher, Evaluations and Benchmarks

ActiveFence

. Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk .

Posted 9/17/2026full-timeRemote • California • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI safety and security evaluations, with a strong background in building and running evaluations for language models. Capable of directing research efforts and collaborating with AI labs while maintaining a robust publication record.

Highest-signal resume keywords
PhD Or Masters In Computer Science3+ Years Building And Running Safety Evaluations5+ Relevant Research Publications In AI SafetyStrong Engineering Skills Including Evaluation HarnessesStrong English Communication Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Safety EvaluationsSecurity EvaluationsEvaluation HarnessesDistributed InferenceTaxonomy DevelopmentSFTDPOGRPOAgentic EvaluationPrompt Injection
Soft Skills
Curiosity About AI HarmsAbility To Direct FreelancersStrong Verbal CommunicationAbility To Present To Large AudiencesWillingness To Travel
Industry Keywords
AI LabsLanguage ModelsResearch PublicationsBenchmark TaxonomyEcosystem Monitoring

About the role

Key responsibilities & impact
  • Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk
  • Collaborate with leading AI labs and universities on benchmarks and papers
  • Own benchmark taxonomy, harness, quality bar, and release
  • Read evals personally and verify taxonomy alignment, rubrics, verifiers, and data distribution
  • Hold the benchmark plan and calendar and keep researchers on timeline
  • Direct two or three freelancers/SMEs as needed
  • Meet roughly monthly with the CTO, pod, and research leads to revise the quarterly release plan
  • Tie the release roadmap to target accounts
  • Spend approximately 20% of time monitoring the ecosystem and reading research
  • Maintain contacts inside AI labs and speak with lab personnel weekly
  • Travel to conferences a couple of times a year

Requirements

What you’ll need
  • PhD or Masters in computer science, machine learning or a related field, or equivalent depth from industry research
  • 3+ years building and running safety or security evaluations for language models in production, at an AI lab, a model provider, or a safety and security research organisation
  • 5+ relevant research publications in AI safety and security, including lead author on at least 2
  • Strong engineering skills, including evaluation harnesses, distributed inference, vLLM, reading a codebase and fixing it
  • Ability to build a taxonomy, not only score against one
  • Ability to direct a researcher and two freelancers without formally managing them
  • Strong English, written and spoken
  • Curiosity about AI harms and ability to learn a new subject every three weeks
  • Ideally: post-training experience with SFT, DPO, or GRPO
  • Ideally: agentic evaluation experience including tool use, orchestration, permissions, and prompt injection
  • Ideally: publications at top conferences
  • Willingness to present own work on client calls
  • Strong verbal and written communication and ability to present to large and/or senior audiences
  • Willingness to travel to conferences at least 3 times a year

Benefits

Comp & perks
  • Budget for freelancers (SMEs) directed ad hoc when needed
  • Travel to conferences at least 3 times a year
  • Opportunity to collaborate with leading AI labs and universities
  • Access to approximately 150 researchers working on harms
  • Conference travel a couple of times a year / at least 3 times a year