FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Data Scientist, AI Evaluations Platform
RBC. Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing evaluation methodologies for generative AI and agentic systems, with a strong focus on metrics, datasets, and quality measurement. Proven ability to lead technical teams and communicate effectively with stakeholders across various domains, ensuring alignment with governance and business objectives.
Highest-signal resume keywords
Evaluation Framework DesignLLM Evaluation MethodsData Science and Machine LearningTechnical Leadership and MentorshipGovernance and Model Risk Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data ScienceMachine LearningEvaluation MethodologiesStatisticsPythonSQLExperimentationData CurationQuality MeasurementMetrics Development
Soft Skills
CommunicationStakeholder ManagementInfluencing Senior LeadersMentorship
Tools & Technologies
MLflowLangfuseLangSmithOpenTelemetryGrafanaCI/CD Pipelines
Industry Keywords
Generative AIAgentic AIModel Risk ManagementResponsible AIFinancial ServicesAudit-Ready Evidence
Tech Stack
Tools & technologiesGrafanaPythonSQL
About the role
Key responsibilities & impact- Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks
- Define evaluation science standards translating model risk, responsible AI, product quality, safety, and business expectations into measurable criteria and repeatable methods
- Own the end-to-end lifecycle for evaluation datasets and scorecards, including sourcing, curation, validation, quality checks, versioning, lineage, reuse, and ongoing improvement
- Design scalable evaluation approaches for generative AI and agentic systems, including task-level, workflow-level, trajectory-level, and runtime evaluation methods
- Partner with AI research, platform engineering, product, risk, governance, and business teams to embed evaluations into build, release, certification, monitoring, and recertification workflows
- Establish human evaluation and review protocols producing reliable labels, reviewer guidance, adjudication processes, quality controls, and audit-ready evidence
- Measure and improve scorer accuracy, calibration, robustness, failure-mode coverage, and explainability across automated and human evaluation approaches
- Provide technical leadership, executive-ready communication, and mentorship to junior data scientists
Requirements
What you’ll need- 8+ years of experience in data science, applied machine learning, AI evaluation, ML quality, or a related technical field, including experience providing technical leadership and mentorship within a team
- Strong experience designing evaluation frameworks for ML, generative AI, or agentic AI systems, including metrics, datasets, benchmarks, rubrics, and quality measurement
- Practical experience with LLM evaluation methods such as LLM-as-judge, deterministic scoring, human evaluation, hallucination assessment, factuality assessment, safety evaluation, or model quality benchmarking
- Strong technical foundation in data science, statistics, machine learning, experimentation, data curation, Python, SQL, and modern AI/ML development practices
- Proven ability to translate governance, model risk, responsible AI, and business requirements into measurable controls, repeatable evaluation processes, and decision-ready evidence
- Strong communication and stakeholder management skills, with the ability to influence senior leaders across research, engineering, product, governance, risk, and business teams
- Experience evaluating agentic AI systems, tool-calling workflows, multi-step reasoning, runtime traces, trajectory scoring, or workflow-level performance
- Experience in financial services, regulated AI, model risk management, responsible AI, enterprise governance, or audit-ready evidence processes
- Familiarity with tools and platforms such as MLflow, Langfuse, LangSmith, OpenTelemetry, Grafana, CI/CD pipelines, or comparable evaluation and observability tooling
- Publications, patents, open-source contributions, or industry work related to AI evaluation, ML quality, AI safety, applied research, or responsible AI
Benefits
Comp & perks- A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
- Leaders who support your development through coaching and managing opportunities
- Ability to make a difference and lasting impact
- Work in a dynamic, collaborative, progressive, and high-performing team
- A world-class training program in financial services
- Opportunities to do challenging work