FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior AI Researcher – Pre-training
Aleph Alpha. Advance the architecture and training of the next generation of foundation models .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing PyTorch-based training workflows for large foundation models, with a strong focus on optimization, scaling laws, and distributed training dynamics. Proven ability to produce high-quality production code and collaborate effectively across teams to enhance model performance and reliability.
Highest-signal resume keywords
Python ProficiencyPyTorch-Based Training WorkflowsMachine Learning ResearchLarge Distributed Training JobsSoftware Engineering Practices
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine LearningOptimizationScaling LawsTransformer Training DynamicsHyperparameter TuningModel Architecture DesignDebugging Distributed SystemsProduction Code DevelopmentExperiment DesignLarge Model Pre-Training
Soft Skills
Effective CollaborationClear Communication
Tools & Technologies
Megatron-LMDeepSpeedTorchtitanGPU ClustersDistributed Frameworks
Industry Keywords
Foundation ModelsTraining RecipesConvergence IssuesCommunication BottlenecksRace ConditionsSynchronization ErrorsSparse Training ApproachesMixture-of-ExpertsOpen-Source ContributionsTop-Tier Publications
Tech Stack
Tools & technologiesDistributed SystemsPythonPyTorch
About the role
Key responsibilities & impact- Advance the architecture and training of the next generation of foundation models
- Own core elements of the training recipe, including optimizers, schedules, and initialization
- Design PyTorch-based architectural improvements to maximize convergence, stability, and training efficiency
- Develop hyperparameter scaling laws and scale-up methodologies
- Use small-scale proxy experiments to predict multi-thousand-GPU behavior and de-risk training decisions
- Investigate convergence issues such as loss spikes and divergence
- Resolve distributed system failures including communication bottlenecks, race conditions, and synchronization errors
- Partner with Compute Performance, Data, Evaluation, and Post-Training teams on system-model co-design
- Translate mathematical reasoning and empirical observations into principled training decisions
- Produce high-quality production code and influence model quality, run reliability, and model efficiency
Requirements
What you’ll need- Proficient in Python and deeply familiar with PyTorch-based training workflows
- Strong track record in machine learning research and software engineering, demonstrated through shipped models, impactful open-source contributions, or published research
- Strong mathematical foundation and comfort reasoning formally about optimisation, scaling behaviour, and training dynamics
- Deep understanding of transformer training dynamics, optimisation, and large distributed training jobs
- Ability to design rigorous experiments and translate empirical observations into robust training decisions
- Hands-on experience pre-training large models (e.g., 7B+ parameters) on substantial infrastructure (e.g., 100+ GPU clusters)
- Strong software engineering practices, including maintainable, well-tested code and reproducible experimentation workflows
- Ability to implement complex model architectures efficiently and reliably and debug issues across model code, training dynamics, and distributed systems
- Effective collaboration and clear communication across teams
- Able to work in Germany and collaborate regularly on site in Heidelberg
- Preferred: experience training LLMs or multimodal models on large GPU clusters using distributed frameworks such as Megatron-LM, DeepSpeed, or torchtitan
- Preferred: familiarity with scaling laws, hyperparameter transfer, or predicting large-scale training behavior from proxy runs
- Preferred: experience profiling distributed jobs and diagnosing training anomalies
- Preferred: exposure to sparse training approaches such as Mixture-of-Experts
- Preferred: top-tier publications, impactful open-source contributions, or significant shipped technical work
Benefits
Comp & perks- 30 days of paid vacation
- Access to a variety of fitness & wellness offerings via Wellhub
- Mental health support through nilo.health
- Substantially subsidized company pension plan
- Subsidized Germany-wide transportation ticket
- Budget for additional technical equipment
- Flexible working hours
- Hybrid working model
- JobRad® Bike Lease