FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI/ML evaluation methodologies, model training, and fine-tuning, with a strong foundation in Python programming. Proven ability to lead technical teams, mentor members, and translate complex findings into actionable insights.
Highest-signal resume keywords
AI/ML Evaluation MethodologiesModel Training and Fine-TuningPython ProgrammingTechnical Team LeadershipAdversarial Machine Learning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI Evaluation FrameworkModel EvaluationAutomated Red TeamingExperiment DesignDataset CreationMetrics AssessmentProduction Software DeliveryAdversarial TestingSafety Fine-TuningInstruction Hardening
Soft Skills
Strong Communication SkillsMentoring
Tools & Technologies
LLMsResearch PlatformsModel Pipeline Integrations
Certifications & Qualifications
BSc in Computer ScienceMSc in Computer Science
Industry Keywords
AI GovernanceScalable Evaluation ServicesAgentic SecurityBenchmark DesignMultimodal Evaluation
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Lead and develop a multidisciplinary team of engineers and researchers
- Own the AI Evaluation Framework roadmap, architecture, delivery, and production reliability
- Advance evaluation methodologies, adversarial testing, and automated red teaming across models and agentic applications
- Translate governance and product requirements into measurable assessments, scoring, and actionable reports
- Drive evaluation-to-hardening feedback loops, including training data generation, model fine-tuning, and system prompt improvements
- Partner with product and model teams to integrate evaluation and hardening into development and release workflows
Requirements
What you’ll need- 7+ years of experience in AI/ML, security research, or software engineering
- 2+ years leading technical teams
- BSc or MSc in computer science or a related field, or equivalent practical experience
- Hands-on experience with LLMs, model evaluation, and model training or fine-tuning
- Strong Python skills
- Experience delivering production software or research platforms
- Experience designing experiments, datasets, and metrics to assess model behavior and validate improvements
- Ability to own technical roadmaps, mentor team members, and deliver across research and engineering
- Strong communication skills and ability to translate complex findings into clear priorities
- Preferred: experience with adversarial ML, AI red teaming, prompt injection, jailbreaks, agentic security, LLM-as-judge, benchmark design, automated adversarial testing, multimodal evaluation, adversarial training, safety fine-tuning, preference optimization, instruction hardening, AI governance, scalable evaluation services, model pipeline integrations, or agentic workflow testing
Benefits
Comp & perks- Opportunities for career advancement and personal growth
- Access to a diverse range of training programs to sharpen your skills
- Performance-based rewards that celebrate your achievements
- Hybrid setup: 3 days in the office and 2 days working from home
