Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Site Reliability Engineer, DGX Cloud

NVIDIA

. Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting .

Posted 9/15/2026full-timeRemote • California • United StatesSenior💰 $168,000 - $333,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expert-level knowledge in Kubernetes administration, containerization, and microservices architecture, with a strong focus on operational reliability and performance at scale. Proficient in building observability stacks and managing GPU workloads across multiple cloud platforms.

Highest-signal resume keywords
Kubernetes AdministrationInfrastructure Automation ToolsSRE PrinciplesObservability Stack DevelopmentGPU Workload Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesContainerizationMicroservices ArchitecturePythonGoLinux Operating SystemsNetworking FundamentalsSLOs/SLIsObservability ToolsAI Inference Workloads
Soft Skills
Incident ManagementRoot-Cause AnalysisCollaboration
Tools & Technologies
TerraformAnsiblePrometheusGrafanaELK StackOpenTelemetryTemporalAirflowKubeVirtCUDA
Certifications & Qualifications
BS in Computer Science
Industry Keywords
Cloud Security StandardsOperational ReliabilityPerformance MonitoringCapacity ManagementBlameless Postmortems

Tech Stack

Tools & technologies
AirflowAnsibleAWSAzureChefCloudGoogle Cloud PlatformGrafanaKubernetesLinuxMicroservicesPrometheusPuppetPythonPyTorchSplunkTCP/IPTerraformGo

About the role

Key responsibilities & impact
  • Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting
  • Support services before launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews
  • Maintain live services by measuring and supervising availability, latency, and overall system health
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
  • Scale systems sustainably through automation and drive changes that improve reliability and velocity
  • Lead triage and root-cause analysis of high-severity incidents
  • Practice balanced incident response and blameless postmortems
  • Participate in on-call rotation to support production services

Requirements

What you’ll need
  • BS in Computer Science or related technical field, or equivalent experience
  • 8+ years of experience operating production services
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
  • Proficiency in at least one high-level programming language (e.g., Python, Go)
  • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management
  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
  • Operating GPU-accelerated clusters with KubeVirt in production
  • Applying generative-AI techniques to reduce operational toil
  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
  • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score