FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and managing large-scale GPU clusters for AI workloads, with a strong foundation in Kubernetes and systems programming languages like Go or Python. Proven ability to collaborate across multifunctional teams to ensure high reliability and performance of AI infrastructure.
Highest-signal resume keywords
Kubernetes API DevelopmentGPU Resource SchedulingCluster Management SystemsSoftware Engineering PrinciplesOperational Excellence
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software DevelopmentSystems Programming LanguageData StructuresAlgorithmsIncident Management Process
Soft Skills
Strong Communication SkillsCollaboration
Tools & Technologies
KubernetesSlurmBright Cluster Manager
Industry Keywords
AI WorkloadsProduction SystemsGPU ClustersTelemetryReliability
Tech Stack
Tools & technologiesCloudKubernetesPythonGo
About the role
Key responsibilities & impact- Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads
- Develop custom software for scheduling GPU resources on Kubernetes
- Implement monitoring and health-management capabilities for reliability, availability, and scalability of GPU assets
- Harness data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry
- Collaborate with teams across NVIDIA to ensure production AI clusters run reliably, consistently, and at maximum performance
- Evaluate system failures and improve services through a defined incident-management process
Requirements
What you’ll need- Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work
- Software development experience with Kubernetes APIs and frameworks, not just operating a cluster
- Strong communication skills and ability to work with multifunctional teams, principals, architects, and across organizational boundaries and geographies
- 5+ years in a similar role and experience on large-scale production systems
- Experience with common software engineering principles, tools, and techniques
- BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree, or equivalent experience
- Technical knowledge of a systems programming language such as Go or Python
- Solid understanding of data structures and algorithms
- Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, or Bright Cluster Manager
- Proven operational excellence maintaining reliable and performant AI infrastructure
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score
