Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Production Engineer – DGX Cloud

NVIDIA

. Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads across DGX Cloud environments .

Posted 10/2/2026full-timeRemote • California • United StatesSenior💰 $184,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in building and operating production services and large-scale distributed systems, with a strong focus on automation, reliability, and incident response. Proficient in using infrastructure as code and GitOps methodologies to enhance service deployment and management.

Highest-signal resume keywords
Production Services ManagementInfrastructure As CodeStrong Programming Skills In PythonSRE Principles UnderstandingAutomation For Service Deployments

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonGoLinuxKubernetesCloud InfrastructureDistributed SystemsNetworking FundamentalsAutomation ToolsConfiguration ManagementService Instrumentation
Soft Skills
Clear Technical CommunicationCollaboration Across Teams
Tools & Technologies
NVIDIA Cloud FunctionsSGLangVLLMNVIDIA DynamoGitOps
Certifications & Qualifications
BS/MS In Computer Science
Industry Keywords
Control Plane ServicesModel DeploymentsInference WorkloadsService HealthSLIsSLOsError BudgetsIncident Response

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesLinuxPythonGo

About the role

Key responsibilities & impact
  • Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads across DGX Cloud environments
  • Improve the reliability of inference and agentic platforms and services, including NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, and inference services built with NVIDIA Dynamo
  • Improve endpoint availability, inference routing, capacity management, and service health
  • Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently across environments
  • Build workflows for service enablement, model releases, handoff, deprecation, and ongoing operations
  • Define and instrument SLIs and SLOs for inference and control plane services and use error budgets to guide reliability improvements
  • Participate in on-call and incident response, troubleshoot failures, and turn recurring issues into automation and durable fixes
  • Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams to design and operate services safely at scale

Requirements

What you’ll need
  • 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation
  • Strong programming skills in Python, Go, or a comparable language
  • Experience developing tools for production operations
  • Experience with infrastructure as code, configuration management, or GitOps
  • Experience building automation for repeatable service deployments and changes
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals
  • Ability to diagnose failures in production
  • Understanding of SRE principles, including SLIs, SLOs, error budgets, incident response, and reducing operational toil
  • Experience instrumenting services and using metrics, logs, and traces to understand system behavior and improve reliability
  • Clear technical communication and ability to work across engineering teams
  • BS/MS in Computer Science or equivalent experience

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score