Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
RBC

Staff Data Scientist, AI Evaluations Platform

RBC

. Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks .

Posted 9/21/2026full-timeToronto • CanadaLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and implementing evaluation methodologies for generative AI and agentic systems, with a strong focus on metrics, datasets, and quality measurement. Proven ability to lead technical teams and communicate effectively with stakeholders across various domains, ensuring alignment with governance and business objectives.

Highest-signal resume keywords
Evaluation Framework DesignLLM Evaluation MethodsData Science and Machine LearningTechnical Leadership and MentorshipGovernance and Model Risk Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data ScienceMachine LearningEvaluation MethodologiesStatisticsPythonSQLExperimentationData CurationQuality MeasurementMetrics Development
Soft Skills
CommunicationStakeholder ManagementInfluencing Senior LeadersMentorship
Tools & Technologies
MLflowLangfuseLangSmithOpenTelemetryGrafanaCI/CD Pipelines
Industry Keywords
Generative AIAgentic AIModel Risk ManagementResponsible AIFinancial ServicesAudit-Ready Evidence

Tech Stack

Tools & technologies
GrafanaPythonSQL

About the role

Key responsibilities & impact
  • Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks
  • Define evaluation science standards translating model risk, responsible AI, product quality, safety, and business expectations into measurable criteria and repeatable methods
  • Own the end-to-end lifecycle for evaluation datasets and scorecards, including sourcing, curation, validation, quality checks, versioning, lineage, reuse, and ongoing improvement
  • Design scalable evaluation approaches for generative AI and agentic systems, including task-level, workflow-level, trajectory-level, and runtime evaluation methods
  • Partner with AI research, platform engineering, product, risk, governance, and business teams to embed evaluations into build, release, certification, monitoring, and recertification workflows
  • Establish human evaluation and review protocols producing reliable labels, reviewer guidance, adjudication processes, quality controls, and audit-ready evidence
  • Measure and improve scorer accuracy, calibration, robustness, failure-mode coverage, and explainability across automated and human evaluation approaches
  • Provide technical leadership, executive-ready communication, and mentorship to junior data scientists

Requirements

What you’ll need
  • 8+ years of experience in data science, applied machine learning, AI evaluation, ML quality, or a related technical field, including experience providing technical leadership and mentorship within a team
  • Strong experience designing evaluation frameworks for ML, generative AI, or agentic AI systems, including metrics, datasets, benchmarks, rubrics, and quality measurement
  • Practical experience with LLM evaluation methods such as LLM-as-judge, deterministic scoring, human evaluation, hallucination assessment, factuality assessment, safety evaluation, or model quality benchmarking
  • Strong technical foundation in data science, statistics, machine learning, experimentation, data curation, Python, SQL, and modern AI/ML development practices
  • Proven ability to translate governance, model risk, responsible AI, and business requirements into measurable controls, repeatable evaluation processes, and decision-ready evidence
  • Strong communication and stakeholder management skills, with the ability to influence senior leaders across research, engineering, product, governance, risk, and business teams
  • Experience evaluating agentic AI systems, tool-calling workflows, multi-step reasoning, runtime traces, trajectory scoring, or workflow-level performance
  • Experience in financial services, regulated AI, model risk management, responsible AI, enterprise governance, or audit-ready evidence processes
  • Familiarity with tools and platforms such as MLflow, Langfuse, LangSmith, OpenTelemetry, Grafana, CI/CD pipelines, or comparable evaluation and observability tooling
  • Publications, patents, open-source contributions, or industry work related to AI evaluation, ML quality, AI safety, applied research, or responsible AI

Benefits

Comp & perks
  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • A world-class training program in financial services
  • Opportunities to do challenging work