Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Research Scientist / Engineer – Training Infrastructure

Luma AI

. Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs .

Posted 10/8/2026full-timeLondon • United KingdomMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and optimizing distributed training systems for large-scale models, with a strong focus on GPU utilization and parallelization techniques. Proficient in building monitoring tools and ensuring the reliability of training runs across extensive GPU clusters.

Highest-signal resume keywords
Distributed PyTorch TrainingGPU Cluster ManagementParallelization TechniquesLinux Systems AdministrationContainerization and Orchestration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed Training SystemsParallelizationGPU OptimizationScriptingMonitoring Tools
Tools & Technologies
NCCLMPICloud Infrastructure
Industry Keywords
Foundation-Model TrainingLarge-Scale TrainingResource UtilizationNetworkingStorage Systems

Tech Stack

Tools & technologies
CloudDistributed SystemsLinuxPyTorch

About the role

Key responsibilities & impact
  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs
  • Research and implement advanced parallelization, including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel
  • Build monitoring, visualization, and debugging tools for large-scale training runs
  • Optimize training stability, convergence, and resource utilization across massive clusters
  • Learn the current training stack and diagnose stability and utilization issues at scale during the first 30 days
  • Ship and validate a parallelization or stability improvement that measurably helps a real training run during days 30–60
  • Build monitoring and tooling to keep large runs reliable and efficient during days 60–90
  • Build distributed systems that train Luma's large-scale multimodal models across thousands of GPUs
  • Enable researchers to focus on innovation on top of reliable, efficient, scalable infrastructure

Requirements

What you’ll need
  • Extensive distributed PyTorch training and parallelisms in foundation-model training
  • Deep understanding of GPU clusters, networking, and storage systems
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization
  • Strong Linux systems administration and scripting
  • Experience managing training runs across 100+ GPUs
  • Experience with containerization, orchestration, and cloud infrastructure

Benefits

Comp & perks
  • Equal opportunity employer
  • Optional diversity and inclusion survey; participation is voluntary and refusal will not affect the job application