FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing GPU infrastructure for AI workloads, with a strong focus on performance validation, automation frameworks, and troubleshooting complex distributed systems. Proven ability to mentor engineers and drive cross-functional initiatives to enhance system reliability and efficiency.
Highest-signal resume keywords
GPU Infrastructure ExpertiseLinux Systems ProficiencyPython Programming SkillsAutomation Frameworks ExperiencePerformance Validation Methodologies
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GPU InfrastructureHigh-Performance Computing (HPC)Distributed SystemsPerformance DebuggingValidation Standards DefinitionCluster ProvisioningPerformance BaselinesAutomation FrameworksNCCL Communication LibrariesSystem Reliability Improvement
Soft Skills
Excellent CommunicationCross-Team CollaborationMentoring
Tools & Technologies
AnsibleLinux SystemsServer HardwareHigh-Speed Networking
Industry Keywords
GPU ClustersPerformance BottlenecksValidation FrameworksTest SuitesProactive TestingTuning
Tech Stack
Tools & technologiesAnsibleDistributed SystemsLinuxPython
About the role
Key responsibilities & impact- Design, optimize, and scale GPU infrastructure for AI training and inference workloads
- Own end-to-end validation of GPU clusters and new hardware platforms
- Lead hardware qualification and bring-up for new GPU platforms
- Analyze and resolve performance bottlenecks across GPU, CPU, PCIe, and network layers
- Establish and maintain validation frameworks, test suites, and performance baselines
- Develop and enhance automation frameworks for cluster provisioning and validation
- Define performance baselines and validation methodologies
- Troubleshoot complex distributed system issues, including communication libraries such as NCCL
- Improve system reliability through proactive testing and tuning
- Mentor engineers and elevate team capabilities
- Drive cross-functional initiatives to improve GPU cluster reliability and efficiency
Requirements
What you’ll need- 5+ years of experience in GPU infrastructure, HPC, or distributed systems
- Strong expertise in Linux systems and server hardware
- Proven experience with large-scale GPU clusters
- Strong programming skills in Python (beyond basic scripting)
- Experience with automation frameworks (Ansible or similar)
- Experience defining validation standards and performance baselines for GPU infrastructure
- Strong debugging skills across system layers (hardware → OS → network)
- Familiarity with high-speed networking concepts
- Excellent communication and cross-team collaboration skills
- Located in the United States
- Legally authorized to work in the United States
- Must answer whether employment visa sponsorship is required
Benefits
Comp & perks- 100% company-paid insurance premiums for employee medical, dental and vision plans
- 401(k) plan that matches 100% up to 4%, with immediate vesting
- Professional Development Reimbursement of $2,500 each year
- 11 Holidays + Paid Time Off Accrual + Rollover Plan
- Increased PTO at 3 year and 10 year anniversary
- 1 month paid sabbatical every 5 years
- Anniversary Bonus each year
- $500 stipend for remote office setup in first year + $400 each following year
- Internet reimbursement up to $75 per month
- Gym membership reimbursement up to $50 per month
- Company paid Wellable subscription
