Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Software Engineer, Distributed Systems Engineering

NVIDIA

. Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads .

Posted 10/9/2026full-timeRemote • California • United StatesSenior💰 $152,000 - $287,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and managing large-scale GPU clusters for AI workloads, with a strong foundation in Kubernetes and systems programming languages like Go or Python. Proven ability to collaborate across multifunctional teams to ensure high reliability and performance of AI infrastructure.

Highest-signal resume keywords
Kubernetes API DevelopmentGPU Resource SchedulingCluster Management SystemsSoftware Engineering PrinciplesOperational Excellence

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software DevelopmentSystems Programming LanguageData StructuresAlgorithmsIncident Management Process
Soft Skills
Strong Communication SkillsCollaboration
Tools & Technologies
KubernetesSlurmBright Cluster Manager
Industry Keywords
AI WorkloadsProduction SystemsGPU ClustersTelemetryReliability

Tech Stack

Tools & technologies
CloudKubernetesPythonGo

About the role

Key responsibilities & impact
  • Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads
  • Develop custom software for scheduling GPU resources on Kubernetes
  • Implement monitoring and health-management capabilities for reliability, availability, and scalability of GPU assets
  • Harness data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry
  • Collaborate with teams across NVIDIA to ensure production AI clusters run reliably, consistently, and at maximum performance
  • Evaluate system failures and improve services through a defined incident-management process

Requirements

What you’ll need
  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work
  • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster
  • Strong communication skills and ability to work with multifunctional teams, principals, architects, and across organizational boundaries and geographies
  • 5+ years in a similar role and experience on large-scale production systems
  • Experience with common software engineering principles, tools, and techniques
  • BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree, or equivalent experience
  • Technical knowledge of a systems programming language such as Go or Python
  • Solid understanding of data structures and algorithms
  • Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, or Bright Cluster Manager
  • Proven operational excellence maintaining reliable and performant AI infrastructure

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score