FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expert-level proficiency in Kubernetes internals and GPU orchestration, with a strong foundation in software engineering and platform management. Capable of leading technical vision, mentoring teams, and driving infrastructure decisions in a multi-tenant environment.
Highest-signal resume keywords
Kubernetes Internals ExpertiseGPU Orchestration in KubernetesTechnical Leadership and MentoringInfrastructure-as-Code and GitOpsObservability at Scale
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Go ProgrammingPython ProgrammingKubernetesDistributed Systems PrinciplesLinux Systems and NetworkingGPU-Aware SchedulingManaged Services DesignChaos EngineeringObservability ToolsHPC and Job Schedulers
Soft Skills
Technical VisionCollaborationMentoringInfluencing Infrastructure DirectionCustomer Engagement
Tools & Technologies
NVIDIA GPU OperatorPrometheusGrafanaDCGMMIGRDMAInfiniBandRoCESlurmCNCF Projects
Industry Keywords
Managed KubernetesMulti-Tenant PlatformsAI WorkloadsIncident-Response AutomationSecurity and ComplianceCapacity PlanningAnomaly DetectionPredictive MaintenanceOpen-Source Community EngagementTechnical Content Creation
Tech Stack
Tools & technologiesDistributed SystemsGrafanaKubernetesLinuxPrometheusPythonGo
About the role
Key responsibilities & impact- Drive technical vision for Lambda's Managed Kubernetes bare-metal platform
- Integrate and extend NVIDIA's open-source ecosystem
- Design GPU-aware orchestration systems
- Lead development of services powering managed services
- Define networking requirements for AI workloads, including CNI, high-performance fabrics, RDMA, and GPUDirect
- Inform storage architecture requirements for AI workloads
- Build the foundation for Managed Slurm on Kubernetes
- Design inference platform services, autoscaling, and multi-model deployment patterns
- Design self-healing systems and incident-response automation
- Lead chaos engineering efforts
- Establish operational excellence through upgrade automation, security patching, and zero-downtime maintenance
- Bridge Orchestration with Network, Storage, and Security teams
- Drive infrastructure-wide decisions and standardization
- Work with customers and internal teams on migration to the managed platform
- Set technical direction, influence roadmap and prioritization, and lead design reviews
- Mentor and grow engineers
- Collaborate with Network, Storage, Security, and Customer Success teams
- Engage with NVIDIA and the open-source community
- Represent Lambda through technical content, conference talks, and customer engagements
- Shape AIOps capabilities for capacity planning, anomaly detection, and predictive maintenance
Requirements
What you’ll need- 10+ years of experience in software engineering, platform engineering, or SRE, with at least 5 years focused on Kubernetes at scale
- Expert-level understanding of Kubernetes internals: API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns
- Holistic infrastructure expertise across compute, networking, storage, and security
- Strong software engineering skills in Go (required) and Python
- Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling
- Proven track record of technical leadership, driving design decisions, mentoring engineers, and influencing infrastructure direction
- Deep experience designing and operating managed services or multi-tenant platforms
- Strong understanding of distributed systems principles: consensus, fault tolerance, consistency models, and graceful degradation
- Experience with observability at scale: Prometheus, Grafana, distributed tracing, and actionable alerting systems
- Solid knowledge of Linux systems and networking (L2-L7), including RDMA, InfiniBand, and RoCE
- Experience with infrastructure-as-code and GitOps workflows
- Experience building and operating managed Kubernetes services or Kubernetes control plane components
- Hands-on experience with NVIDIA's open-source ecosystem beyond GPU Operator
- Familiarity with HPC and traditional job schedulers such as Slurm and Kubernetes-native batch scheduling
- Background in confidential computing
- Experience migrating customers or workloads from legacy/bespoke infrastructure to standardized platforms
- Contributions to CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects
- Familiarity with security and compliance in multi-tenant environments
- Background in ML infrastructure, including training clusters, inference serving, or simulation
- Ability and willingness to work onsite at the San Francisco office 4 days a week
- Legal authorization to work in the United States
Benefits
Comp & perks- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan
- AI-assisted development tools
- World-class team and opportunities to work with NVIDIA's GPU and networking stack
- Technical blog posts, conference talks, and strategic customer engagement opportunities
