Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
ActiveFence

Researcher, Evaluations and Benchmarks

ActiveFence

. Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk .

Posted 9/17/2026full-timeRemote • New York • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI safety and security evaluations, with a strong background in building taxonomies and managing research collaborations. Proven ability to lead projects, communicate effectively, and maintain high-quality standards in research outputs.

Highest-signal resume keywords
PhD Or Masters In Computer Science3+ Years In Safety Or Security Evaluations5+ Relevant Research Publications In AI SafetyStrong Engineering Skills Including Evaluation HarnessesStrong English Communication Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Safety EvaluationsSecurity EvaluationsEvaluation HarnessesDistributed InferenceTaxonomy DevelopmentSFT Post-Training ExperienceDPO Post-Training ExperienceGRPO Post-Training ExperienceAgentic Evaluation ExperienceResearch Publication
Soft Skills
Curiosity About HarmsAbility To Direct ResearchersStrong Verbal CommunicationAbility To Present To Large Audiences
Industry Keywords
AI SafetyAI SecurityLanguage ModelsFreelance ManagementBenchmarkingResearch CollaborationQuality AssuranceClient CommunicationResearch Ecosystem MonitoringConference Presentations

About the role

Key responsibilities & impact
  • Ship a benchmark every two to three weeks measuring a previously unmeasured frontier risk
  • Collaborate with leading AI labs and universities on benchmarks and papers
  • Pair with in-house researchers who own specific harm areas and direct freelancers
  • Own the taxonomy, evaluation harness, quality bar, and release
  • Ensure frontier labs can rerun the benchmark and reproduce the reported numbers
  • Read evals, verify taxonomy alignment, and push researchers on quality
  • Own the benchmark plan and calendar and keep researchers on schedule
  • Direct two or three freelance subject matter experts as needed
  • Meet roughly monthly with the CTO, pod, and research leads to set the quarterly release roadmap
  • Tie the release plan to target accounts
  • Monitor the AI safety and security ecosystem, read research, maintain lab contacts, and gather feedback
  • Travel to conferences and speak with AI lab contacts weekly

Requirements

What you’ll need
  • PhD or Masters in computer science, machine learning or a related field, or equivalent depth from industry research
  • 3+ years building and running safety or security evaluations for language models in production, at an AI lab, a model provider, or a safety and security research organisation
  • 5+ relevant research publications in AI safety and security, including lead author on at least 2
  • Strong engineering skills, including evaluation harnesses, distributed inference, vLLM, reading a codebase and fixing it
  • Ability to build a taxonomy, not only score against one
  • Ability to direct a researcher and two freelancers without formally managing them
  • Strong English, written and spoken
  • Curiosity about harms and ability to learn a new subject every three weeks
  • Post-training experience with SFT, DPO, or GRPO (ideally)
  • Reward design for subjective and safety-relevant targets (ideally)
  • Agentic evaluation experience with tool use, orchestration, permissions, and prompt injection (ideally)
  • Publications at top conferences (ideally)
  • Willingness to present work on client calls; strong verbal and written communication and ability to present to large and/or senior audiences (ideally)

Benefits

Comp & perks
  • Budget for freelancers directed ad hoc when needed
  • Travel to conferences a couple of times a year
  • Travel to conferences at least 3 times a year (ideally)