Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Site Reliability Engineer, Production Engineering

NVIDIA

. Lead a global, dynamic, state-of-the-art Service Reliability Operations center .

Posted 9/20/2026full-timeRemote • IndiaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in administering large-scale production Kubernetes systems, with a strong focus on automation, incident management, and service reliability. Proficient in Linux system administration and familiar with high-performance computing environments.

Highest-signal resume keywords
Kubernetes AdministrationService Reliability OperationsLinux System AdministrationCI/CD Tools ExperienceIncident Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesSLURMLarge-Scale Cluster ManagementScriptingPythonGolangRust
Soft Skills
Strong Communication SkillsPersuasive Presentation
Tools & Technologies
JenkinsArgoCDObservability Tools
Industry Keywords
Cloud ProductsHigh-AvailabilityData CenterGPU HardwareDPU HardwareCore Linux Networking

Tech Stack

Tools & technologies
CloudDNSFirewallsJenkinsKubernetesLinuxPythonRustGo

About the role

Key responsibilities & impact
  • Lead a global, dynamic, state-of-the-art Service Reliability Operations center
  • Provide support for NVIDIA Cloud products and services
  • Partner with Site Reliability Engineering, Security Operations Center, DevOps teams, and other organizations
  • Support Production Kubernetes Services with a focus on automation and reducing manual tasks
  • Perform large-scale Kubernetes administration, systems administration, and security monitoring to maintain service SLAs, integrity, and reliability
  • Use alerts, alarms, and observability tools to monitor, detect, prevent, and respond to incidents
  • Analyze logs, metrics, and system behavior to troubleshoot issues
  • Lead root cause analysis and implement effective resolutions
  • Initiate and lead incident management calls
  • Coordinate subject matter experts and service owners for timely incident escalation and resolution
  • Develop monitors, alarms, and alerts to improve service reliability and customer experience

Requirements

What you’ll need
  • 7+ years of demonstrated experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments
  • Strong preference for on-prem expertise
  • BS in Computer Science, Engineering, Mathematics, or equivalent experience
  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
  • Familiarity with GPU / DPU hardware and high-performance computing Cluster environments
  • Strong Linux system administration, DNS, DHCP and core Linux networking (IP Tables, routing, firewalls) experience
  • Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure
  • Experience working with CI/CD tools like Jenkins, ArgoCD
  • Experience in scripting
  • Programming in Python or Golang or Rust preferred, but not required
  • Strong communication and soft skills, able to present to cross-functional group members in a persuasive manner

Benefits

Comp & perks
  • 24/7 Production engineering team support
  • Flexibility to work on split-weekend shifts