Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Research Scientist / Engineer – Training Infrastructure

Luma AI

. Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs .

Posted 10/8/2026full-timeRemoteMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and optimizing distributed training systems using PyTorch, with a focus on advanced parallelization techniques and GPU cluster management. Proficient in building monitoring tools and ensuring training stability and efficiency across large-scale environments.

Highest-signal resume keywords
Distributed PyTorch TrainingGPU Cluster ManagementAdvanced Parallelization TechniquesMonitoring and Debugging ToolsCommunication Libraries (NCCL, MPI)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed Training SystemsParallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel)Training Stability OptimizationResource Utilization OptimizationLinux Systems AdministrationScriptingMulti-Node TrainingContainerizationOrchestrationCloud Infrastructure
Industry Keywords
Foundation-Model TrainingLarge-Scale Training RunsNetworkingStorage SystemsDistributed-System Optimization

Tech Stack

Tools & technologies
CloudLinuxNode.jsPyTorch

About the role

Key responsibilities & impact
  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs
  • Research and implement advanced parallelization, including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel
  • Build monitoring, visualization, and debugging tools for large-scale training runs
  • Optimize training stability, convergence, and resource utilization across massive clusters
  • Learn the current training stack and diagnose stability and utilization issues at scale
  • Deliver a parallelization or stability improvement that measurably benefits a real training run
  • Build monitoring and tooling to keep large runs reliable and efficient

Requirements

What you’ll need
  • Extensive distributed PyTorch training and parallelisms in foundation-model training
  • Deep understanding of GPU clusters, networking, and storage systems
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization
  • Strong Linux systems administration and scripting (nice to have)
  • Experience managing training runs across 100+ GPUs (nice to have)
  • Experience with containerization, orchestration, and cloud infrastructure (nice to have)
  • Experience at the level of FSDP and multi-node training
  • Ability to work remotely in the EU

Benefits

Comp & perks
  • Equal opportunity employer
  • Voluntary diversity and inclusion survey participation; refusal will not affect the job application
  • Remote work arrangement