FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Software Engineer, SRE, Production Engineering – DGX Cloud
NVIDIA. Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in building automation for bare-metal provisioning and managing NVIDIA GPU hardware, with strong skills in Go or Python for production infrastructure. Capable of diagnosing failures across complex systems and collaborating effectively across teams to implement durable solutions.
Highest-signal resume keywords
Go ProgrammingPython ProgrammingBMC and Redfish ExperienceNVIDIA GPU Hardware ManagementLinux and Kubernetes Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Bare-Metal ProvisioningProduction AutomationFirmware ManagementNetwork BootIncident HandlingDebugging Hardware FailuresServer Lifecycle ManagementDPU Mode OperationWorkflow AutomationCluster Performance Validation
Soft Skills
Clear CommunicationOwnership of Problems
Tools & Technologies
NVIDIA NVL72 SystemsBlueField-3 DPUsGitOpsArgo CDSLOs
Certifications & Qualifications
BS/MS in Computer Science
Industry Keywords
Cloud Partner EnvironmentsData Center OperationsObservabilityRoot-Cause AnalysisHardware Validation
Tech Stack
Tools & technologiesCloudKubernetesLinuxPythonGo
About the role
Key responsibilities & impact- Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management
- Build tools using BMC and Redfish interfaces to assess hardware health, regulate server state, and facilitate recovery workflows
- Manage and enhance NVIDIA NVL72 systems and BlueField-3 or later DPUs within cloud partner and on-premises environments
- Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues into automated detection and repair
- Define validation and handoff criteria so new capacity enters production safely and consistently
- Take part in on-call duties, incident response, root-cause analysis, and ensure permanent resolutions are implemented
- Collaborate with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries
Requirements
What you’ll need- 8+ years building software for or operating production infrastructure, including substantial hands-on bare-metal experience
- Strong Go or Python skills, with a record of delivering production automation and services
- Direct experience working with BMC and Redfish for server provisioning, health inspection, power control, or fault diagnosis
- Practical experience working directly with NVIDIA GPU hardware, including NVL72 systems, and BlueField-3 or later DPUs
- Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair
- Experience managing production reliability via on-call duties, incident handling, observability, and durable solutions
- Ability to debug failures across hardware, host operating systems, networking, and distributed services
- Clear communication and demonstrated ownership of problems that span multiple teams
- BS/MS in Computer Science or equivalent experience in a related field
- Experience operating BlueField DPUs in DPU mode, including host-to-DPU connectivity and lifecycle debugging, or equivalent experience
- Background with NVLink, InfiniBand, Spectrum-X, or GPU cluster performance validation
- Experience building safe, repeatable workflows for rack-scale bringup, firmware upgrades, hardware replacement, and customer handoff
- Background with Kubernetes, GitOps, Argo CD, SLOs, and fleet-wide automation
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score