Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Staff AI Training Infrastructure Engineer

Designworks Talent LLC

. Build and scale distributed training infrastructure supporting large AI models across large GPU clusters .

Posted 10/2/2026full-timeBellevue • Washington • United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and scaling distributed training infrastructure for large AI models, with a strong focus on reliability, efficiency, and resource utilization. Proficient in integrating AI models into production pipelines and optimizing GPU performance within large-scale environments.

Highest-signal resume keywords
Distributed Training SystemsLarge-Scale Machine Learning InfrastructurePyTorch DistributedKubernetesGPU Utilization Optimization

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Programming SkillsFault ToleranceCheckpointingRecovery SolutionsSupervised Fine-TuningReinforcement Learning from Human FeedbackMulti-Node GPU TrainingPerformance EngineeringAutomation ToolsComplex Distributed Systems
Soft Skills
Problem SolvingProject Ownership
Tools & Technologies
DeepSpeedMegatron-LMRayContainerized AI WorkloadsAI Infrastructure Platforms
Industry Keywords
AI ModelsTraining InfrastructureOperational ProcessesProduction Machine Learning PipelinesHyperscaler

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesNode.jsPyTorchRay

About the role

Key responsibilities & impact
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters
  • Design and improve systems that increase training reliability, efficiency, and resource utilization
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations
  • Integrate AI models into production training pipelines with platform, orchestration, and performance engineering teams
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency
  • Build tools and automation that improve the developer experience for AI researchers and engineers
  • Establish best practices for training infrastructure, operational processes, and platform reliability
  • Contribute to the evolution of the AI infrastructure platform as an early engineering team member

Requirements

What you’ll need
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems
  • Strong understanding of reliability, scalability, and efficiency challenges associated with multi-node GPU training
  • Experience integrating training systems with production machine learning pipelines
  • Strong programming skills and experience with complex distributed systems
  • Ability to independently own technically challenging projects
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
  • Experience optimizing GPU utilization, training performance, or distributed system reliability
  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms
  • U.S. work authorization required
  • Visa sponsorship is not currently available

Benefits

Comp & perks
  • Certain roles are eligible for merit increases, annual bonus, and long term incentives based on individual performance
  • Medical, dental, and vision insurance for U.S.-based employees
  • 401(k) plan and company match
  • Paid holidays per calendar year
  • Hybrid work arrangement