FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expert-level backend development proficiency in Python, Java/Kotlin, TypeScript/JavaScript, or Go, with a strong focus on creating evaluations for customer-facing agents. Proven ability to lead and mentor engineering teams while collaborating with data science peers to develop proprietary benchmarks and establish industry-accepted evaluation methods.
Highest-signal resume keywords
Backend Development ProficiencyTechnical LeadershipAgent Evaluation ExperienceIntegration Onboarding ProcessesCollaboration with Data Science
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonJavaKotlinTypeScriptJavaScriptGoFastAPINode.jsAgent Evaluation TechniquesTask Completion Measurement
Soft Skills
MentoringCollaborationLeadership
Tools & Technologies
OpenAIAnthropicGoogle LLMClaude CodeCodexOpencodePiPlaywrightChrome DevToolsMCP Servers
Certifications & Qualifications
BS/BA Degree in Related Field
Industry Keywords
Agent-Oriented ProductsEvaluation FrameworksBenchmark FrameworksDurable ExecutionWorkflow Frameworks
Tech Stack
Tools & technologiesJavaJavaScriptKotlinNode.jsPythonTypeScriptGo
About the role
Key responsibilities & impact- Scope feasibility and effort for evaluating software-vendor agents
- Own the full stack of the core evaluation system, including admin and API surfaces and evaluation workflow dispatch
- Design reusable evaluation system primitives across software verticals
- Turn integration onboarding and maintenance processes into repeatable AI skills or agents
- Monitor emerging agent-evaluation frameworks and techniques
- Establish industry-accepted methods for distributing evaluation results
- Collaborate with data science peers to develop proprietary benchmarks
- Lead, mentor, and provide technical leadership to a core engineering team
- Promote the use of evaluations in agent-oriented products across the organization
- Deliver useful and credible evaluations of live software-vendor agents
- Apply agent-first engineering techniques to deliver work to production
Requirements
What you’ll need- 10+ years of professional programming experience in backend or full-stack environments
- 2+ years of experience directly managing engineers
- Expert-level backend development proficiency using Python, Java/Kotlin, TypeScript/JavaScript, or Go
- Strong proficiency with backend frameworks such as FastAPI or Node.js
- Direct experience creating evaluations against customer-facing agents using agent trajectory trace data and rubrics
- Experience measuring task completion rate, accuracy, correctness, or policy adherence
- Direct experience using frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications
- Regular use of coding agent harnesses such as Claude Code, Codex, Opencode, or Pi
- BS/BA degree in a related field
- Experience with agent tool use via direct integration or MCP servers is advantageous
- Knowledge of STATE-Bench, tau2-bench, or similar benchmark frameworks is advantageous
- Experience with Playwright, browser-use, Chrome DevTools MCP, or similar tools is advantageous
- Experience with durable execution/workflow frameworks or agent sandboxing is advantageous
- Authorization to work in the United States without restriction
- No current or future U.S. work sponsorship requirement
Benefits
Comp & perks- Equity
- Bonus
- Flexible work
- Ample parental leave
- Unlimited PTO
- Inclusive and diverse work environment
- Employee resource groups (ERGs)
- G2 Gives philanthropic program
- AI-assisted application review opt-out option
