FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing machine learning systems, particularly in molecular ML and cheminformatics, with a strong focus on data validation, reproducibility, and production-grade model delivery in regulated environments. Proficient in translating scientific requirements into actionable pipeline behaviors while ensuring data integrity and performance reporting.
Highest-signal resume keywords
Machine Learning Systems DevelopmentMolecular ML or CheminformaticsPython ProgrammingML Ops or ML InfrastructureProduction-Grade Model Delivery
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine LearningPythonData ValidationModel EvaluationKubernetesRDKitPerformance ReportingData Leakage PreventionVersion ControlModular Code Development
Soft Skills
CollaborationProblem-SolvingCommunication
Tools & Technologies
PyTorchKubernetesChEMBLTox21ToxCast
Industry Keywords
ADMETToxicity PredictionBiotechPharmaceuticalFederated Learning
Tech Stack
Tools & technologiesKubernetesPythonPyTorch
About the role
Key responsibilities & impact- Own model pipelines end to end, from partner data landing through data preparation, training, evaluation and release of benchmarked model weights for ADMET and toxicity endpoints
- Build schema and data-contract validation, actionable-error validators, and QC/profiling reports usable without access to partners' raw data
- Make federated runs reproducible and auditable through versioned configurations, pinned data snapshots and model provenance
- Build leakage-safe splitting, held-out benchmarks and honest performance reporting
- Turn prototypes and research code into tested, modular systems and hand them to engineering for scaling into Foundry
- Translate scientific requirements into pipeline behavior with the science team
- Surface data or modeling risks early to partners and internal teams
Requirements
What you’ll need- 5+ years building ML systems in Python, with version control, tested modular code, code review, and interfaces other people can use
- Hands-on molecular ML or cheminformatics using RDKit, fingerprints and descriptors, or graph/transformer models, applied to property, activity or toxicity prediction
- Experience building training and evaluation pipelines that other people run, not one-off notebooks
- Understanding of molecular ML failure modes including data leakage, split design, applicability domain, dataset shift, and over-optimistic benchmarks
- Ability to write validators and data contracts under partial visibility, where raw data cannot be inspected
- Comfortable with PyTorch or an equivalent modern ML stack
- Ability to translate ambiguous scientific requirements into defined, testable pipeline behavior
- Federated learning, privacy-preserving ML, or other multi-party training environments
- ML Ops or ML infrastructure experience, particularly Kubernetes-based training, evaluation or deployment workflows
- Production-grade model delivery in regulated, enterprise, pharmaceutical or biotech settings
- Familiarity with public ADMET, toxicity and bioactivity data resources including ChEMBL, Tox21 and ToxCast
- Open-source contributions or a publication record in cheminformatics, molecular ML, or applied machine learning
Benefits
Comp & perks- Industry-competitive compensation, including early-stage virtual share options
- Remote-first working - work where you work best
- Wellbeing budget
- Mental health support
- Work-from-home budget
- Co-working stipend
- Learning budget
- Generous holiday allowance
- Office days at our Berlin HQ or a different European location (3x per year)
