Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Evolent

Senior Data Scientist – AI Evaluation, Improvement

Evolent

. Own the diagnostic loop for LLM-based clinical services .

Posted 10/8/2026full-timeRemote • United StatesSenior💰 $135,000 - $165,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in evaluating and improving LLM-based clinical services through structured error analysis, experiment design, and metrics interpretation. Proficient in handling clinical data under security and compliance requirements while mentoring team members and collaborating with clinical reviewers.

Highest-signal resume keywords
LLM Evaluation MethodsData Science ExperiencePython ProgrammingExperiment DesignClinical Data Handling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data ScienceApplied Machine LearningLLM EvaluationStructured Error AnalysisMetrics InterpretationExperiment DesignPrompt ConfigurationData AnalysisSQLHypothesis-Driven Working Style
Soft Skills
Strong Written CommunicationMentoring
Tools & Technologies
LLM-as-JudgeTracing ToolsObservability ToolingRegression SuitesGolden Sets
Certifications & Qualifications
Bachelor's Degree in Data ScienceMaster's Degree in Quantitative Field (Preferred)
Industry Keywords
Clinical DataPHI ComplianceHealthcare ExperienceExperimentation Background

Tech Stack

Tools & technologies
PythonSQL

About the role

Key responsibilities & impact
  • Own the diagnostic loop for LLM-based clinical services
  • Form root-cause hypotheses for failure modes involving prompts, context, retrieval, guidelines, model behavior, and upstream data
  • Design experiments to isolate causes
  • Test prompt, configuration, context, and model variants
  • Verify improvements through structured evaluations and confirm no regressions
  • Own the evaluation roadmap and quality bar for AI services
  • Build golden sets and regression suites
  • Define and interpret metrics including accuracy, guideline adherence, and grounding/faithfulness
  • Extend the team's evaluation platform thoughtfully
  • Partner with clinical reviewers to turn findings into labeled evidence and evaluation criteria
  • Partner with product managers to prioritize important failure modes
  • Mentor others and help grow the evaluations function
  • Support the annual clinical-guidelines update cycle with regression evaluation
  • Document failure taxonomies, experiment templates, and variant history
  • Handle clinical data, including PHI, under security, privacy, and compliance requirements

Requirements

What you’ll need
  • Bachelor's degree in Data Science, Computer Science, Statistics, or a related quantitative field — or equivalent experience
  • 5+ years of data science or applied machine learning experience, including shipping and maintaining models or AI systems in production
  • 2+ years of recent, hands-on experience evaluating and improving LLM-based systems
  • Experience with structured error analysis, prompt/configuration iteration, experiment design, and metrics interpretation
  • Experience with LLM evaluation methods and tools, including golden/regression sets, LLM-as-judge with validation, and tracing/observability tooling
  • Strong Python and solid data-analysis skills; SQL is a plus
  • Comfort computing and reasoning about sensitivity, specificity, and PPV
  • Hypothesis-driven working style
  • Strong written communication
  • Must handle clinical data, including PHI, according to security, privacy, and compliance requirements
  • Comprehensive background check required
  • In-person I-9 verification required
  • May be subject to drug screening prior to employment
  • Government-issued photo ID required for identity verification
  • High-speed internet over 10 Mbps at home
  • Preferred: healthcare experience; mentoring or technical workstream leadership; experience working with clinical reviewers or domain experts; LLM observability/telemetry stacks; statistics or experimentation background; Master's degree in a quantitative field

Benefits

Comp & perks
  • Work/life balance
  • Flexible work arrangements and autonomy
  • Comprehensive benefits, including health insurance benefits, for qualifying employees
  • High-speed home internet requirement supported as a home-work capability