Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Principal Software Engineer, Distributed Systems Engineer – DGX Cloud

NVIDIA

. Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads .

Posted 9/30/2026full-timeRemote • North Carolina • United StatesLead💰 $272,000 - $431,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in software engineering with a focus on Kubernetes for managing large-scale GPU clusters, ensuring reliability and performance through effective monitoring and incident management. Proficient in systems programming languages such as Go or Python, with a solid understanding of data structures and algorithms.

Highest-signal resume keywords
Kubernetes Cluster OperationsGPU Resource SchedulingSoftware Development with Kubernetes APIsSystems Programming in Go or PythonOperational Excellence in AI Infrastructure

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringKubernetesCluster Management SystemsData StructuresAlgorithmsIncident ManagementMonitoring and Health ManagementDistributed Systems AutomationGPU Hardware DiagnosticsTelemetry Analysis
Soft Skills
Strong Communication SkillsCollaboration with Multifunctional Teams
Tools & Technologies
Kubernetes APIsSlurmBright Cluster Manager
Certifications & Qualifications
BS in Computer ScienceBS in EngineeringBS in PhysicsBS in Mathematics
Industry Keywords
AI WorkloadsLarge-Scale Production SystemsTechnical OrganizationOperational Excellence

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesNode.jsPythonGo

About the role

Key responsibilities & impact
  • Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads
  • Develop custom software for scheduling GPU resources on Kubernetes
  • Implement monitoring and health-management capabilities for reliability, availability, and scalability of GPU assets
  • Harness data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry
  • Collaborate with teams across NVIDIA to ensure production AI clusters run reliably, consistently, and with maximum performance
  • Evaluate system failures and improve services through a defined incident-management process

Requirements

What you’ll need
  • Significant software engineering experience with Kubernetes, including cluster operations, operator development, node health monitoring, and GPU resource scheduling
  • Direct experience in a software engineering role within a highly technical organization, with demonstrable impact
  • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster
  • Strong communication skills and ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies
  • 15+ years in a similar role and experience with large-scale production systems
  • Experience with common software engineering principles, tools, and techniques
  • BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree, or equivalent experience
  • Systems programming language proficiency, including Go or Python
  • Solid understanding of data structures and algorithms
  • Technical competency managing and automating large-scale distributed systems independent of cloud providers
  • Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, or Bright Cluster Manager
  • Proven operational excellence maintaining reliable and performant AI infrastructure

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score