FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating AI evaluation platforms, utilizing methodologies that integrate human and AI expert judgment. Proficient in Python and familiar with machine learning frameworks, capable of producing insightful reports for stakeholders.
Highest-signal resume keywords
Python ProgrammingMachine Learning IntegrationAI Model EvaluationCross-Functional CollaborationTechnical Mentorship
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software Platform DevelopmentAutomated MetricsEvaluation MethodologiesAPI FrameworksComputer VisionLanguage ModelingTask-Specific Agent EvaluationPyTorchHugging Face EcosystemExperiment Provenance
Soft Skills
Strong Communication SkillsLeadership PotentialCross-Functional Collaboration Experience
Certifications & Qualifications
Active TS/SCI Clearance
Industry Keywords
AI PlatformBenchmarkingEvaluation HarnessesOperational BaselinesUser Interface DesignAgentic WorkflowsFailure ModesStakeholder ReportingKnowledge TransferHybrid Work Environment
Tech Stack
Tools & technologiesPythonPyTorchTypeScript
About the role
Key responsibilities & impact- Build and operate benchmarking and evaluation capabilities at the core of an AI platform
- Build evaluation harnesses using automated metrics and structured human expert judgment
- Apply evaluations to candidate models and agentic workflows
- Develop repeatable methodologies for comparing performance against operational baselines
- Document limitations and surface important failure modes
- Produce reports for senior stakeholders to inform capability fielding decisions
- Design and implement a platform for testing and evaluating AI models and agentic systems
- Develop evaluation methodologies combining human and AI expert judging across multiple input and output modalities
- Design and enhance intuitive, generic workflow user interfaces
- Collaborate with subject matter experts, model/agent developers, and evaluation designers
- Architect experiment provenance and result-tracking systems
- Ensure performance across integrated tools and scaling systems
- Provide technical mentorship, code reviews, design guidance, documentation, knowledge transfer, and user onboarding
Requirements
What you’ll need- U.S. citizenship required
- U.S.-remote candidates must be willing and able to undergo the process required to obtain and maintain a U.S. security clearance
- Hybrid candidates must be located in or willing to work hybrid from Washington, DC; Denver, CO; or Colorado Springs, CO
- Hybrid opening requires an active TS/SCI clearance
- Proven experience building and deploying software platforms with complex integration surfaces, preferably handling machine learning or agentic workloads
- Deep experience using Python, including API frameworks and agentic harnesses
- High-level knowledge of computer vision, language modeling, and/or task-specific agent evaluation
- Hands-on experience with PyTorch and the Hugging Face ecosystem
- Strong communication skills and cross-functional collaboration experience, including external stakeholders
- Demonstrated leadership potential or experience in architectural decisions, coordinating engineers, and developing standards and best practices
- Up to 15% travel required for the hybrid position
Benefits
Comp & perks- 100% employer paid medical premiums for employees
- Self-managed PTO with a minimum time off requirement
- Career framework providing a pathway and recognition of increased impact
- Opportunities to collaborate with global experts
- Open-source project participation
- Learning, innovation, and ownership support
- Flexible remote or hybrid work arrangements depending on position
