Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Senior AI Training Infrastructure Engineer

Designworks Talent LLC

. Build and scale distributed training infrastructure supporting large AI models across large GPU clusters .

Posted 10/3/2026full-timeBellevue • Washington • United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and scaling distributed training infrastructure for large AI models, with a strong focus on reliability, efficiency, and resource utilization. Proficient in integrating AI models into production pipelines and optimizing performance within complex distributed systems.

Highest-signal resume keywords
Distributed Training SystemsLarge-Scale Machine Learning InfrastructureGPU Utilization OptimizationAI Model IntegrationProgramming Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed SystemsFault ToleranceCheckpointingRecovery SolutionsMulti-Node GPU TrainingSupervised Fine-TuningReinforcement LearningPerformance EngineeringAutomation ToolsOperational Processes
Soft Skills
Problem SolvingCollaborationOwnershipIndependenceAdaptability
Tools & Technologies
PyTorch DistributedDeepSpeedMegatron-LMRayKubernetes
Industry Keywords
AI InfrastructureMachine Learning PipelinesHyperscalerCloud ProviderGPU Cloud Environment

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesNode.jsPyTorchRay

About the role

Key responsibilities & impact
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters
  • Design and improve systems that increase training reliability, efficiency, and resource utilization
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations
  • Integrate AI models into production training pipelines with platform, orchestration, and performance engineering teams
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency
  • Build tools and automation that improve the developer experience for AI researchers and engineers
  • Establish best practices for training infrastructure, operational processes, and platform reliability
  • Contribute to the evolution of the AI infrastructure platform as an early engineering team member
  • Collaborate with infrastructure, orchestration, performance, and machine learning teams

Requirements

What you’ll need
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems
  • Strong understanding of reliability, scalability, and efficiency challenges associated with multi-node GPU training
  • Experience integrating training systems with production machine learning pipelines
  • Strong programming skills and experience working with complex distributed systems
  • Ability to independently own technically challenging projects
  • Comfortable operating with high ownership and limited process overhead
  • Preferred: Experience with PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
  • Preferred: Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
  • Preferred: Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
  • Preferred: Experience optimizing GPU utilization, training performance, or distributed system reliability
  • Preferred: Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms
  • U.S. work authorization is required
  • Visa sponsorship is not currently available

Benefits

Comp & perks
  • Certain roles are eligible for merit increases
  • Annual bonus eligibility for certain roles
  • Long-term incentives for certain roles
  • Medical insurance for U.S.-based employees
  • Dental insurance for U.S.-based employees
  • Vision insurance for U.S.-based employees
  • 401(k) plan
  • Company 401(k) match
  • Paid holidays per calendar year