FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Scientist – GenAI Evaluation
Johnson & Johnson. Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining evaluation frameworks for AI/ML systems, with a strong focus on scientific accuracy and regulatory compliance. Proficient in Python and modern AI/ML tools, with the ability to translate scientific judgment into measurable evaluation criteria.
Highest-signal resume keywords
AI/ML EvaluationPython ProficiencyGenerative AI ExperienceEvaluation Framework DesignRegulatory Environment Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI/ML EvaluationData ScienceEvaluation Framework DesignPrompt EngineeringScientific BenchmarkingQuality AssessmentCI/CD Evaluation PipelinesLLM-as-Judge SystemsData CurationStatistical Analysis
Soft Skills
CommunicationCollaborationAnalytical ThinkingProblem Solving
Tools & Technologies
LLM APIsVector DatabasesEmbedding ModelsEvaluation HarnessesAutomated Evaluation Pipelines
Certifications & Qualifications
Master's Degree in AI/ML or Related FieldPhD Preferred
Industry Keywords
Biomedical EngineeringDrug Development PipelineOncologyImmunologyNeuroscienceFDAEMABiomedical InformaticsSynthetic DatasetsHuman Expert Agreement
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy
- Author evaluation rubrics and scoring criteria
- Curate golden and synthetic datasets with domain experts
- Maintain a registry of reusable evaluation assets
- Validate AI judges against human expert agreement
- Run model, prompt, retriever, and agent benchmarks producing standardized quality readouts
- Analyze failure patterns such as hallucination, unsupported claims, and weak traceability
- Turn evaluation findings into actionable recommendations
- Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners
- Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality
- Build evaluation tooling and reusable patterns enabling other teams to self-serve
- Inform AI release decisions for pharmaceutical R&D workflows
Requirements
What you’ll need- Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field required
- 6+ years of hands-on experience in AI/ML evaluation or data science for candidates with a Master's degree, or 3+ years of industry experience for candidates with a PhD
- Experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems
- Hands-on experience with generative AI, including large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering
- Strong proficiency in Python and modern AI/ML tooling, including evaluation harnesses, embedding models, vector databases, and LLM APIs
- Ability to translate expert scientific judgment into measurable criteria, rubrics, and reproducible protocols
- English proficiency required, written and verbal
- PhD preferred
- Experience with AI/ML evaluation in regulated environments such as FDA, EMA, or equivalent preferred
- Understanding of the drug development pipeline and biomedical data types preferred
- Domain expertise in oncology, immunology, or neuroscience preferred
- Experience designing or validating LLM-as-judge systems preferred
- Experience implementing CI/CD evaluation pipelines preferred
- Publications in AI evaluation, NLP, or biomedical informatics preferred
Benefits
Comp & perks- Annual bonus with set target depending on pay grade/location, with actual amount based on employee and company performance
- Vacation days
- Parental leave for a minimum of 12 weeks
- Bereavement leave
- Caregiver leave
- Volunteer leave
- Well-being reimbursement
- Programs for financial, physical, and mental health
- Service anniversary and recognition awards
- Insurance plans for employees and, in some locations, eligible dependents