FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in leading Runtime Engineering teams, managing complex Kubernetes configurations, and ensuring operational reliability across distributed systems. Proficient in cluster lifecycle management, security standards, and effective communication with cross-functional teams.
Highest-signal resume keywords
Kubernetes Internals KnowledgeCluster Lifecycle ManagementPeople-Management ExperienceNVIDIA GPU Operator ExperienceOpen Source Contributions
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cluster APIKubeadmCIS Kubernetes BenchmarkAPI Design and ImplementationContainer Runtime Stack ManagementCluster Networking (CNI)Storage Management (CSI)Pod Security PoliciesMulti-Tenancy IsolationGPU Resource Partitioning
Soft Skills
Effective CommunicationCoaching SkillsProblem SolvingTechnical PlanningCross-Organization Consensus Building
Tools & Technologies
NVIDIA Kubernetes Engine (NKE)AICRDCGMIdentity and Access ManagementGPU Node Pools
Certifications & Qualifications
BS/MS Degree in Computer Science or Related Field
Industry Keywords
Distributed Software SystemsProduction-Critical SoftwareRuntime EcosystemHyperscale KubernetesSecurity Standards
Tech Stack
Tools & technologiesCyber SecurityKubernetesNode.jsOpen Source
About the role
Key responsibilities & impact- Lead the Runtime Engineering team responsible for the full configuration lifecycle of NVIDIA Kubernetes Engine (NKE) tenant workload clusters
- Build, implement, and ensure operational reliability of cluster configurations for NKE tenant workloads across supported topologies
- Manage engineers coordinating the container runtime stack, including AICR, GPU management operator, DCGM, and related node-level components
- Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA, and GPU resource partitioning (MIG, MPS, time-slicing)
- Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries
- Partner with NKE platform, infrastructure, and cybersecurity teams to integrate capabilities and resolve runtime concerns
- Build and maintain tooling for AICR lifecycle management, including provisioning, upgrades, configuration drift detection, and remediation
- Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership
- Contribute to open source communities where NKE has upstream dependencies or influence
Requirements
What you’ll need- BS/MS degree in Computer Science or related field (or equivalent experience)
- 12+ overall years of relevant experience designing and delivering large-scale distributed software systems
- 5+ years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software
- Experience leading engineers with varying specializations and seniority levels across runtime, networking, and security
- Kubernetes internals knowledge, including scheduler, kubelet, API server, and admission controllers
- Cluster lifecycle management experience with Cluster API, kubeadm, or equivalent
- Experience leading fleet-scale cluster provisioning and upgrades
- Knowledge of CIS Kubernetes Benchmark, pod security admission, image signing, and supply chain integrity
- Ability to design and implement maintainable APIs for consumers
- Familiarity with Identity and Access Management approaches
- Ability to manage up, down, and across organizations
- Ability to reach cross-organization consensus
- Prior experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling
- Experience running Kubernetes at hyperscale with GPU node pools
- Track record of upstream open source contributions in the Kubernetes or open source runtime ecosystem
- Effective written, verbal, and engineering-leadership communication skills
- Skills in coaching, analysis, problem solving, and short/long-term technical planning
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package
- Equity
- Benefits for you and your family
