Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Johnson & Johnson

Senior Scientist – GenAI Evaluation

Johnson & Johnson

. Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy .

Posted 9/24/2026full-timeMadrid • SpainSenior💰 €55,400 - €87,860 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and maintaining evaluation frameworks for AI/ML systems, with a strong focus on scientific accuracy and regulatory compliance. Proficient in Python and modern AI/ML tools, with the ability to translate scientific judgment into measurable evaluation criteria.

Highest-signal resume keywords
AI/ML EvaluationPython ProficiencyGenerative AI ExperienceEvaluation Framework DesignRegulatory Environment Experience

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI/ML EvaluationData ScienceEvaluation Framework DesignPrompt EngineeringScientific BenchmarkingQuality AssessmentCI/CD Evaluation PipelinesLLM-as-Judge SystemsData CurationStatistical Analysis
Soft Skills
CommunicationCollaborationAnalytical ThinkingProblem Solving
Tools & Technologies
LLM APIsVector DatabasesEmbedding ModelsEvaluation HarnessesAutomated Evaluation Pipelines
Certifications & Qualifications
Master's Degree in AI/ML or Related FieldPhD Preferred
Industry Keywords
Biomedical EngineeringDrug Development PipelineOncologyImmunologyNeuroscienceFDAEMABiomedical InformaticsSynthetic DatasetsHuman Expert Agreement

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy
  • Author evaluation rubrics and scoring criteria
  • Curate golden and synthetic datasets with domain experts
  • Maintain a registry of reusable evaluation assets
  • Validate AI judges against human expert agreement
  • Run model, prompt, retriever, and agent benchmarks producing standardized quality readouts
  • Analyze failure patterns such as hallucination, unsupported claims, and weak traceability
  • Turn evaluation findings into actionable recommendations
  • Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners
  • Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality
  • Build evaluation tooling and reusable patterns enabling other teams to self-serve
  • Inform AI release decisions for pharmaceutical R&D workflows

Requirements

What you’ll need
  • Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field required
  • 6+ years of hands-on experience in AI/ML evaluation or data science for candidates with a Master's degree, or 3+ years of industry experience for candidates with a PhD
  • Experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems
  • Hands-on experience with generative AI, including large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering
  • Strong proficiency in Python and modern AI/ML tooling, including evaluation harnesses, embedding models, vector databases, and LLM APIs
  • Ability to translate expert scientific judgment into measurable criteria, rubrics, and reproducible protocols
  • English proficiency required, written and verbal
  • PhD preferred
  • Experience with AI/ML evaluation in regulated environments such as FDA, EMA, or equivalent preferred
  • Understanding of the drug development pipeline and biomedical data types preferred
  • Domain expertise in oncology, immunology, or neuroscience preferred
  • Experience designing or validating LLM-as-judge systems preferred
  • Experience implementing CI/CD evaluation pipelines preferred
  • Publications in AI evaluation, NLP, or biomedical informatics preferred

Benefits

Comp & perks
  • Annual bonus with set target depending on pay grade/location, with actual amount based on employee and company performance
  • Vacation days
  • Parental leave for a minimum of 12 weeks
  • Bereavement leave
  • Caregiver leave
  • Volunteer leave
  • Well-being reimbursement
  • Programs for financial, physical, and mental health
  • Service anniversary and recognition awards
  • Insurance plans for employees and, in some locations, eligible dependents