Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Lambda

Staff Software Engineer – Managed Kubernetes

Lambda

. Drive technical vision for Lambda's Managed Kubernetes bare-metal platform .

Posted 9/25/2026full-timeUnited StatesLead💰 $314,000 - $465,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expert-level proficiency in Kubernetes internals and GPU orchestration, with a strong foundation in software engineering and platform management. Capable of leading technical vision, mentoring teams, and driving infrastructure decisions in a multi-tenant environment.

Highest-signal resume keywords
Kubernetes Internals ExpertiseGPU Orchestration in KubernetesTechnical Leadership and MentoringInfrastructure-as-Code and GitOpsObservability at Scale

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Go ProgrammingPython ProgrammingKubernetesDistributed Systems PrinciplesLinux Systems and NetworkingGPU-Aware SchedulingManaged Services DesignChaos EngineeringObservability ToolsHPC and Job Schedulers
Soft Skills
Technical VisionCollaborationMentoringInfluencing Infrastructure DirectionCustomer Engagement
Tools & Technologies
NVIDIA GPU OperatorPrometheusGrafanaDCGMMIGRDMAInfiniBandRoCESlurmCNCF Projects
Industry Keywords
Managed KubernetesMulti-Tenant PlatformsAI WorkloadsIncident-Response AutomationSecurity and ComplianceCapacity PlanningAnomaly DetectionPredictive MaintenanceOpen-Source Community EngagementTechnical Content Creation

Tech Stack

Tools & technologies
Distributed SystemsGrafanaKubernetesLinuxPrometheusPythonGo

About the role

Key responsibilities & impact
  • Drive technical vision for Lambda's Managed Kubernetes bare-metal platform
  • Integrate and extend NVIDIA's open-source ecosystem
  • Design GPU-aware orchestration systems
  • Lead development of services powering managed services
  • Define networking requirements for AI workloads, including CNI, high-performance fabrics, RDMA, and GPUDirect
  • Inform storage architecture requirements for AI workloads
  • Build the foundation for Managed Slurm on Kubernetes
  • Design inference platform services, autoscaling, and multi-model deployment patterns
  • Design self-healing systems and incident-response automation
  • Lead chaos engineering efforts
  • Establish operational excellence through upgrade automation, security patching, and zero-downtime maintenance
  • Bridge Orchestration with Network, Storage, and Security teams
  • Drive infrastructure-wide decisions and standardization
  • Work with customers and internal teams on migration to the managed platform
  • Set technical direction, influence roadmap and prioritization, and lead design reviews
  • Mentor and grow engineers
  • Collaborate with Network, Storage, Security, and Customer Success teams
  • Engage with NVIDIA and the open-source community
  • Represent Lambda through technical content, conference talks, and customer engagements
  • Shape AIOps capabilities for capacity planning, anomaly detection, and predictive maintenance

Requirements

What you’ll need
  • 10+ years of experience in software engineering, platform engineering, or SRE, with at least 5 years focused on Kubernetes at scale
  • Expert-level understanding of Kubernetes internals: API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns
  • Holistic infrastructure expertise across compute, networking, storage, and security
  • Strong software engineering skills in Go (required) and Python
  • Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling
  • Proven track record of technical leadership, driving design decisions, mentoring engineers, and influencing infrastructure direction
  • Deep experience designing and operating managed services or multi-tenant platforms
  • Strong understanding of distributed systems principles: consensus, fault tolerance, consistency models, and graceful degradation
  • Experience with observability at scale: Prometheus, Grafana, distributed tracing, and actionable alerting systems
  • Solid knowledge of Linux systems and networking (L2-L7), including RDMA, InfiniBand, and RoCE
  • Experience with infrastructure-as-code and GitOps workflows
  • Experience building and operating managed Kubernetes services or Kubernetes control plane components
  • Hands-on experience with NVIDIA's open-source ecosystem beyond GPU Operator
  • Familiarity with HPC and traditional job schedulers such as Slurm and Kubernetes-native batch scheduling
  • Background in confidential computing
  • Experience migrating customers or workloads from legacy/bespoke infrastructure to standardized platforms
  • Contributions to CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects
  • Familiarity with security and compliance in multi-tenant environments
  • Background in ML infrastructure, including training clusters, inference serving, or simulation
  • Ability and willingness to work onsite at the San Francisco office 4 days a week
  • Legal authorization to work in the United States

Benefits

Comp & perks
  • Generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan
  • AI-assisted development tools
  • World-class team and opportunities to work with NVIDIA's GPU and networking stack
  • Technical blog posts, conference talks, and strategic customer engagement opportunities