Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA

. Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments .

Posted 9/23/2026full-timeRemote • California • United StatesSenior💰 $184,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production infrastructure, with a strong focus on automation for GPU clusters and cloud environments. Proficient in troubleshooting distributed systems and improving operational workflows through advanced tools and methodologies.

Highest-signal resume keywords
Python ProgrammingKubernetes ExperienceInfrastructure AutomationGitOps ImplementationIncident Response

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production InfrastructureDistributed Systems TroubleshootingCloud InfrastructureLinux ProficiencyGPU InfrastructureKubernetes OperatorsTerraformArgoCDFleet AutomationSLOs
Soft Skills
Clear CommunicationTeam Collaboration
Tools & Technologies
APIsGitOpsAutomation ToolsMonitoring ToolsIncident Response Tools
Certifications & Qualifications
BS/MS in Computer Science
Industry Keywords
BMaaSVMaaSManaged KubernetesMulti-Cloud InfrastructureObservability Practices

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesLinuxPythonTerraformGo

About the role

Key responsibilities & impact
  • Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments
  • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations
  • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows
  • Participate in on-call, incident response, debugging, and durable follow-up work
  • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready

Requirements

What you’ll need
  • 8+ years of experience building or operating production infrastructure
  • Strong programming skills in Python, Go, or similar
  • Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation
  • Ability to troubleshoot distributed systems in production
  • Clear communication and ability to work across teams
  • BS/MS in Computer Science or equivalent experience
  • Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation
  • Experience with SLOs, on-call, incident response, observability, and reliability practices
  • Exposure to BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score