Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
FAR.AI

Research Scientist, Applied White-Box Methods

FAR.AI

. Take ownership of and accelerate the Applied White-Box Methods team's research agenda .

Posted 9/23/2026full-timeBerkeley • California • United StatesMid-LevelSenior💰 $150,000 - $250,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI safety and white-box methods, with a strong ability to develop and evaluate innovative techniques for improving AI systems. Engages effectively with both technical and non-technical audiences while contributing to the AI alignment community.

Highest-signal resume keywords
White-Box Method ApplicationAI Safety ResearchEvaluations and AI ControlPhD in Computer Science or Related FieldApplied Machine Learning Experience

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Activation ExplainersSteeringAttributionInfluence FunctionsReinforcement LearningPost-Training of LLMsGoodhart-Resistant MetricsLong-Context Agentic CodingResearch TamperingReward Hacking
Soft Skills
Clear CommunicationCollaboration
Industry Keywords
AI AlignmentRed-TeamingFrontier-Scale DeploymentLarge-Scale ModelsTechnical Writing

About the role

Key responsibilities & impact
  • Take ownership of and accelerate the Applied White-Box Methods team's research agenda
  • Develop, evaluate, and demonstrate methods leveraging model internals to improve AI safety
  • Stress-test white-box methods on real-world tasks, including long-context agentic coding
  • Study applications including white-box control, evaluation awareness, training-dynamics shaping, alignment, reward hacking, sandbagging, and research tampering
  • Define realistic evaluations with Goodhart-resistant metrics
  • Develop methods using realistic, large-scale models and evaluations
  • Compare white-box methods against black-box methods and activation probes
  • Ensure practical monitors and interventions are simple and efficient enough for frontier-scale deployment
  • Publish findings broadly and engage with the AI alignment community
  • Propose new research directions within the team's agenda
  • Attend relevant conferences and community events
  • Collaborate with national AI safety institutes, frontier model developers, and top academics
  • Complete interviews with technical staff followed by a paid work trial of up to one week

Requirements

What you’ll need
  • Hands-on experience applying at least one white-box method to a real model, such as activation explainers, SAEs, steering, attribution, influence functions, or probes, and an informed view of its limitations
  • Track record in AI safety through a paper, fellowship project, or substantive public writing
  • Experience with evaluations, AI control, red-teaming, reinforcement learning, or post-training of LLMs
  • Ability to communicate novel methods and results clearly to technical and non-technical audiences
  • PhD or several years of research experience in computer science, machine learning, physics, statistics, or a related field
  • Previous experience in applied ML for fields such as biology, chemistry, materials science, or robotics
  • For researchers new to LLM research: ability to explain how previous research experience could be leveraged for this work and how they are engaging with technical AI safety research
  • Willingness to work in person in Berkeley or relocate for the in-person role, unless considered for an exceptional remote arrangement

Benefits

Comp & perks
  • Health insurance: 94% of insurance premium paid by Organization commencing within 1 month after your start date
  • 401(k) plan with up to 2% match
  • 25 days Paid Time Off per year, accrued weekly
  • Up to 10 days of paid sick leave per year
  • Paid Bereavement, Family, Medical and Pregnancy Disability Leave
  • Work computer and WFH stipend provided for eligible employees
  • Catered lunches and dinners on workdays at the Berkeley office
  • Work-related travel and equipment expenses paid
  • Visa sponsorship for in-person employees