FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing distributed training systems using PyTorch, with a focus on advanced parallelization techniques and GPU cluster management. Proficient in building monitoring tools and ensuring training stability and efficiency across large-scale environments.
Highest-signal resume keywords
Distributed PyTorch TrainingGPU Cluster ManagementAdvanced Parallelization TechniquesMonitoring and Debugging ToolsCommunication Libraries (NCCL, MPI)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed Training SystemsParallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel)Training Stability OptimizationResource Utilization OptimizationLinux Systems AdministrationScriptingMulti-Node TrainingContainerizationOrchestrationCloud Infrastructure
Industry Keywords
Foundation-Model TrainingLarge-Scale Training RunsNetworkingStorage SystemsDistributed-System Optimization
Tech Stack
Tools & technologiesCloudLinuxNode.jsPyTorch
About the role
Key responsibilities & impact- Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs
- Research and implement advanced parallelization, including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel
- Build monitoring, visualization, and debugging tools for large-scale training runs
- Optimize training stability, convergence, and resource utilization across massive clusters
- Learn the current training stack and diagnose stability and utilization issues at scale
- Deliver a parallelization or stability improvement that measurably benefits a real training run
- Build monitoring and tooling to keep large runs reliable and efficient
Requirements
What you’ll need- Extensive distributed PyTorch training and parallelisms in foundation-model training
- Deep understanding of GPU clusters, networking, and storage systems
- Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization
- Strong Linux systems administration and scripting (nice to have)
- Experience managing training runs across 100+ GPUs (nice to have)
- Experience with containerization, orchestration, and cloud infrastructure (nice to have)
- Experience at the level of FSDP and multi-node training
- Ability to work remotely in the EU
Benefits
Comp & perks- Equal opportunity employer
- Voluntary diversity and inclusion survey participation; refusal will not affect the job application
- Remote work arrangement
