FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer, SRE, Production Engineering
NVIDIA. Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management .
Posted 9/30/2026full-timeRemote • California • United StatesSenior💰 $152,000 - $287,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building automation for bare-metal provisioning and managing NVIDIA GPU hardware, with strong skills in Go or Python for production infrastructure. Proven ability to diagnose and resolve complex issues across hardware and software environments while ensuring production reliability.
Highest-signal resume keywords
Go ProgrammingPython ProgrammingBMC and Redfish ExperienceNVIDIA GPU Hardware ManagementLinux and Firmware Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Bare-Metal ProvisioningAutomation DevelopmentIncident ResponseRoot-Cause AnalysisServer Lifecycle ManagementNetwork BootDebugging Hardware FailuresObservabilityProduction Reliability ManagementCluster Lifecycle Management
Soft Skills
Clear CommunicationOwnership of Problems
Tools & Technologies
NVIDIA NVL72 SystemsBlueField-3 DPUsKubernetesData Center Operations
Certifications & Qualifications
BS/MS in Computer Science
Industry Keywords
Production InfrastructureHardware ValidationFirmware UpgradesServer ProvisioningHealth InspectionPower ManagementFault Diagnosis
Tech Stack
Tools & technologiesCloudKubernetesLinuxPythonGo
About the role
Key responsibilities & impact- Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management
- Develop tools that interact with BMC and Redfish interfaces to monitor hardware health, manage server state, and assist recovery workflows
- Handle and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs throughout cloud partner and on-premises environments
- Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues into automated detection and repair
- Define validation and handoff criteria so new capacity enters production safely and consistently
- Take part in on-call duties, incident response, root-cause analysis, and follow-up to implement permanent solutions
- Work with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries
Requirements
What you’ll need- 5+ years building software for or operating production infrastructure, including substantial hands-on bare-metal experience
- Strong Go or Python skills, with a record of delivering production automation and services
- Direct experience with BMC and Redfish in server provisioning, health inspection, power management, or fault diagnosis
- Practical experience working directly with NVIDIA GPU hardware, such as NVL72 systems, and BlueField-3 or newer DPUs
- Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair
- Experience managing production reliability through on-call duties, incident response, observability, and durable solutions
- Ability to debug failures across hardware, host operating systems, networking, and distributed services
- Clear communication and demonstrated ownership of problems that span multiple teams
- BS/MS in Computer Science or equivalent experience in a practical setting
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score