FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI evaluation design, including defining metrics and ground truth, while effectively collaborating with cross-functional teams to drive quality standards and actionable insights. Proficient in quantitative measurement and statistical analysis, with a strong foundation in Python and SQL for model evaluation in production environments.
Highest-signal resume keywords
Quantitative Measurement RigorStatistical AnalysisPython ProficiencySQL ProficiencyCross-Functional Collaboration
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model EvaluationMetric ValidationSample SizingConfidence IntervalsSignificance TestingAutomated Grader ValidationLLM EvaluationEvaluation HarnessesText-to-SQL EvaluationAnalytics Agent Evaluation
Soft Skills
Excellent CommunicationStrong Problem-Solving Ability
Tools & Technologies
AI ToolsProduction Environments
Industry Keywords
FintechBrokerageRisk Consequences
Tech Stack
Tools & technologiesPythonSQL
About the role
Key responsibilities & impact- Design AI evaluations by defining ground truth, metrics, and scoring methods for models and agents
- Build repeatable evaluation loops to track quality over time and catch regressions before release
- Partner with engineering and analytics engineering to operationalize evaluation harnesses
- Translate evaluation results into actionable recommendations for system improvements
- Establish evaluation guidelines, documentation, and review practices
- Mentor and align teams on evaluation best practices and measurable AI quality
- Partner with Product, Engineering, Analytics Engineering, and business stakeholders to define quality standards and drive iteration
- Own the quality bar independently of the teams that build and optimize the systems
Requirements
What you’ll need- Track record of quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation)
- Strong statistical and ML foundation, including sample sizing, confidence intervals, significance, handling non-determinism, and validating automated graders against human ground truth
- Proficiency in Python and SQL, with experience evaluating models in production environments
- Strong judgment in defining quality metrics and ground truth for ambiguous outputs
- Excellent communication and cross-functional collaboration skills
- Strong problem-solving ability in fast-paced, greenfield environments
- 6–10 years in quantitative data science or ML, with focused experience in measurement or evaluation
- Quantitative degree is a plus; equivalent industry experience is equally welcome
- Nice to have: hands-on LLM/agent evaluation in production, eval harnesses, LLM-as-judge calibration, and CI regression gates
- Nice to have: experience evaluating text-to-SQL, analytics agents, or systems where correctness is verifiable against data
- Nice to have: background in fintech, brokerage, or domains with business or risk consequences
- Nice to have: fluency with AI tools in research and engineering workflows
Benefits
Comp & perks- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
