FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in administering large-scale production Kubernetes systems, with a strong focus on automation, incident management, and service reliability. Proficient in Linux system administration and familiar with high-performance computing environments.
Highest-signal resume keywords
Kubernetes AdministrationService Reliability OperationsLinux System AdministrationCI/CD Tools ExperienceIncident Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesSLURMLarge-Scale Cluster ManagementScriptingPythonGolangRust
Soft Skills
Strong Communication SkillsPersuasive Presentation
Tools & Technologies
JenkinsArgoCDObservability Tools
Industry Keywords
Cloud ProductsHigh-AvailabilityData CenterGPU HardwareDPU HardwareCore Linux Networking
Tech Stack
Tools & technologiesCloudDNSFirewallsJenkinsKubernetesLinuxPythonRustGo
About the role
Key responsibilities & impact- Lead a global, dynamic, state-of-the-art Service Reliability Operations center
- Provide support for NVIDIA Cloud products and services
- Partner with Site Reliability Engineering, Security Operations Center, DevOps teams, and other organizations
- Support Production Kubernetes Services with a focus on automation and reducing manual tasks
- Perform large-scale Kubernetes administration, systems administration, and security monitoring to maintain service SLAs, integrity, and reliability
- Use alerts, alarms, and observability tools to monitor, detect, prevent, and respond to incidents
- Analyze logs, metrics, and system behavior to troubleshoot issues
- Lead root cause analysis and implement effective resolutions
- Initiate and lead incident management calls
- Coordinate subject matter experts and service owners for timely incident escalation and resolution
- Develop monitors, alarms, and alerts to improve service reliability and customer experience
Requirements
What you’ll need- 7+ years of demonstrated experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments
- Strong preference for on-prem expertise
- BS in Computer Science, Engineering, Mathematics, or equivalent experience
- Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
- Familiarity with GPU / DPU hardware and high-performance computing Cluster environments
- Strong Linux system administration, DNS, DHCP and core Linux networking (IP Tables, routing, firewalls) experience
- Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure
- Experience working with CI/CD tools like Jenkins, ArgoCD
- Experience in scripting
- Programming in Python or Golang or Rust preferred, but not required
- Strong communication and soft skills, able to present to cross-functional group members in a persuasive manner
Benefits
Comp & perks- 24/7 Production engineering team support
- Flexibility to work on split-weekend shifts
