FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer, DGX Cloud
NVIDIA. Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting .
Posted 9/15/2026full-timeRemote • California • United StatesSenior💰 $168,000 - $333,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expert-level knowledge in Kubernetes administration, containerization, and microservices architecture, with a strong focus on operational reliability and performance at scale. Proficient in building observability stacks and managing GPU workloads across multiple cloud platforms.
Highest-signal resume keywords
Kubernetes AdministrationInfrastructure Automation ToolsSRE PrinciplesObservability Stack DevelopmentGPU Workload Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesContainerizationMicroservices ArchitecturePythonGoLinux Operating SystemsNetworking FundamentalsSLOs/SLIsObservability ToolsAI Inference Workloads
Soft Skills
Incident ManagementRoot-Cause AnalysisCollaboration
Tools & Technologies
TerraformAnsiblePrometheusGrafanaELK StackOpenTelemetryTemporalAirflowKubeVirtCUDA
Certifications & Qualifications
BS in Computer Science
Industry Keywords
Cloud Security StandardsOperational ReliabilityPerformance MonitoringCapacity ManagementBlameless Postmortems
Tech Stack
Tools & technologiesAirflowAnsibleAWSAzureChefCloudGoogle Cloud PlatformGrafanaKubernetesLinuxMicroservicesPrometheusPuppetPythonPyTorchSplunkTCP/IPTerraformGo
About the role
Key responsibilities & impact- Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting
- Define SLOs/SLIs, monitor error allowances, and streamline reporting
- Support services before launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews
- Maintain live services by measuring and supervising availability, latency, and overall system health
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
- Scale systems sustainably through automation and drive changes that improve reliability and velocity
- Lead triage and root-cause analysis of high-severity incidents
- Practice balanced incident response and blameless postmortems
- Participate in on-call rotation to support production services
Requirements
What you’ll need- BS in Computer Science or related technical field, or equivalent experience
- 8+ years of experience operating production services
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
- Proficiency in at least one high-level programming language (e.g., Python, Go)
- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
- Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
- Operating GPU-accelerated clusters with KubeVirt in production
- Applying generative-AI techniques to reduce operational toil
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
- Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score