FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI safety and white-box methods, with a strong ability to develop and evaluate innovative techniques for improving AI systems. Engages effectively with both technical and non-technical audiences while contributing to the AI alignment community.
Highest-signal resume keywords
White-Box Method ApplicationAI Safety ResearchEvaluations and AI ControlPhD in Computer Science or Related FieldApplied Machine Learning Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Activation ExplainersSteeringAttributionInfluence FunctionsReinforcement LearningPost-Training of LLMsGoodhart-Resistant MetricsLong-Context Agentic CodingResearch TamperingReward Hacking
Soft Skills
Clear CommunicationCollaboration
Industry Keywords
AI AlignmentRed-TeamingFrontier-Scale DeploymentLarge-Scale ModelsTechnical Writing
About the role
Key responsibilities & impact- Take ownership of and accelerate the Applied White-Box Methods team's research agenda
- Develop, evaluate, and demonstrate methods leveraging model internals to improve AI safety
- Stress-test white-box methods on real-world tasks, including long-context agentic coding
- Study applications including white-box control, evaluation awareness, training-dynamics shaping, alignment, reward hacking, sandbagging, and research tampering
- Define realistic evaluations with Goodhart-resistant metrics
- Develop methods using realistic, large-scale models and evaluations
- Compare white-box methods against black-box methods and activation probes
- Ensure practical monitors and interventions are simple and efficient enough for frontier-scale deployment
- Publish findings broadly and engage with the AI alignment community
- Propose new research directions within the team's agenda
- Attend relevant conferences and community events
- Collaborate with national AI safety institutes, frontier model developers, and top academics
- Complete interviews with technical staff followed by a paid work trial of up to one week
Requirements
What you’ll need- Hands-on experience applying at least one white-box method to a real model, such as activation explainers, SAEs, steering, attribution, influence functions, or probes, and an informed view of its limitations
- Track record in AI safety through a paper, fellowship project, or substantive public writing
- Experience with evaluations, AI control, red-teaming, reinforcement learning, or post-training of LLMs
- Ability to communicate novel methods and results clearly to technical and non-technical audiences
- PhD or several years of research experience in computer science, machine learning, physics, statistics, or a related field
- Previous experience in applied ML for fields such as biology, chemistry, materials science, or robotics
- For researchers new to LLM research: ability to explain how previous research experience could be leveraged for this work and how they are engaging with technical AI safety research
- Willingness to work in person in Berkeley or relocate for the in-person role, unless considered for an exceptional remote arrangement
Benefits
Comp & perks- Health insurance: 94% of insurance premium paid by Organization commencing within 1 month after your start date
- 401(k) plan with up to 2% match
- 25 days Paid Time Off per year, accrued weekly
- Up to 10 days of paid sick leave per year
- Paid Bereavement, Family, Medical and Pregnancy Disability Leave
- Work computer and WFH stipend provided for eligible employees
- Catered lunches and dinners on workdays at the Berkeley office
- Work-related travel and equipment expenses paid
- Visa sponsorship for in-person employees
