FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in deploying and managing AI Compute and HPC infrastructure in Linux environments, with strong skills in scripting, cluster management, and advanced networking. Proven ability to interact with customers and internal teams to deliver effective solutions and support for complex systems.
Highest-signal resume keywords
Linux System AdministrationScripting Proficiency (Bash, Python, Ansible)Cluster Management and ProvisioningExperience with AI Compute ProjectsInterpersonal Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux System AdministrationScripting Proficiency (Bash, Python, Ansible)Cluster ManagementPerformance ReportingNetwork RoutingExperience with SLURM, LSF, UGEBenchmarking Tools (HPL, NCCL, MLPerf)InfiniBand ExperienceGPU Hardware/Software ExperienceMPI (Message Passing Interface)
Soft Skills
Excellent Interpersonal SkillsStrong Organizational SkillsAbility to PrioritizeMulti-tasking Ability
Tools & Technologies
KubernetesLustreGPFSBase Command Manager (BCM)
Certifications & Qualifications
Bachelor's Degree in Computer Science or Engineering
Industry Keywords
AI ComputeHPC InfrastructureNetworkingSystem DesignAutomationValidation
Tech Stack
Tools & technologiesAnsibleKubernetesLinuxPython
About the role
Key responsibilities & impact- Deploy, manage, and validate AI Compute/HPC infrastructure in Linux-based environments for new and existing customers
- Act as the domain expert with customers during planning calls through implementation
- Perform handover-related documentation and knowledge transfers to support customers rolling out sophisticated systems
- Provide feedback to internal teams by opening bugs, documenting workarounds, and suggesting improvements
- Interact with customers, partners, and internal teams to analyze, define, and implement large-scale AI Compute projects involving networking, system design, automation, and validation
Requirements
What you’ll need- 8+ years providing in-depth support and deployment services; solving problems for hardware and software products
- Knowledge and experience with Linux system administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, network-routing/advanced networking (tuning and monitoring)
- Cluster management and provisioning technologies for bare-metal servers (bonus credit for BCM (Base Command Manager))
- Minimum of a four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering or equivalent experience
- Scripting proficiency (Bash, Python, Ansible, etc.)
- Excellent interpersonal skills and the ability to deliver resolutions for customer issues as they arise
- Strong organizational skills and ability to prioritize/multi-task easily with limited supervision
- Experience with schedulers such as SLURM, LSF, UGE, etc.
- An ability to travel to customer sites within the United States up to 20% of the time
- Experience with benchmarking tools such as HPL, NCCL tests, MLPerf as well as Kubernetes experience
- InfiniBand experience
- Experience with GPU (Graphics Processing Unit) focused hardware/software
- Experience with MPI (Message Passing Interface)
- Storage technologies such as Lustre or GPFS
- Familiarity with OEM GPU platforms
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score
