FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior AI Training Infrastructure Engineer
Designworks Talent LLC. Build and scale distributed training infrastructure supporting large AI models across large GPU clusters .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and scaling distributed training infrastructure for large AI models, with a strong focus on reliability, efficiency, and resource utilization. Proficient in integrating AI models into production pipelines and optimizing performance within complex distributed systems.
Highest-signal resume keywords
Distributed Training SystemsLarge-Scale Machine Learning InfrastructureGPU Utilization OptimizationAI Model IntegrationProgramming Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed SystemsFault ToleranceCheckpointingRecovery SolutionsMulti-Node GPU TrainingSupervised Fine-TuningReinforcement LearningPerformance EngineeringAutomation ToolsOperational Processes
Soft Skills
Problem SolvingCollaborationOwnershipIndependenceAdaptability
Tools & Technologies
PyTorch DistributedDeepSpeedMegatron-LMRayKubernetes
Industry Keywords
AI InfrastructureMachine Learning PipelinesHyperscalerCloud ProviderGPU Cloud Environment
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetesNode.jsPyTorchRay
About the role
Key responsibilities & impact- Build and scale distributed training infrastructure supporting large AI models across large GPU clusters
- Design and improve systems that increase training reliability, efficiency, and resource utilization
- Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations
- Integrate AI models into production training pipelines with platform, orchestration, and performance engineering teams
- Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency
- Build tools and automation that improve the developer experience for AI researchers and engineers
- Establish best practices for training infrastructure, operational processes, and platform reliability
- Contribute to the evolution of the AI infrastructure platform as an early engineering team member
- Collaborate with infrastructure, orchestration, performance, and machine learning teams
Requirements
What you’ll need- Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
- Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems
- Strong understanding of reliability, scalability, and efficiency challenges associated with multi-node GPU training
- Experience integrating training systems with production machine learning pipelines
- Strong programming skills and experience working with complex distributed systems
- Ability to independently own technically challenging projects
- Comfortable operating with high ownership and limited process overhead
- Preferred: Experience with PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
- Preferred: Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
- Preferred: Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
- Preferred: Experience optimizing GPU utilization, training performance, or distributed system reliability
- Preferred: Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms
- U.S. work authorization is required
- Visa sponsorship is not currently available
Benefits
Comp & perks- Certain roles are eligible for merit increases
- Annual bonus eligibility for certain roles
- Long-term incentives for certain roles
- Medical insurance for U.S.-based employees
- Dental insurance for U.S.-based employees
- Vision insurance for U.S.-based employees
- 401(k) plan
- Company 401(k) match
- Paid holidays per calendar year