FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining scalable Kubernetes control plane services, with strong programming skills in Go and Python. Proficient in building resilient distributed systems and managing GPU-aware orchestration for AI/ML workloads.
Highest-signal resume keywords
Kubernetes InternalsGo ProgrammingPython ProgrammingDistributed Systems FundamentalsObservability at Scale
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesGoPythonDistributed SystemsControl Plane ManagementGPU SchedulingLinux SystemsNetworkingCloud InfrastructureObservability Tools
Tools & Technologies
PrometheusGrafanaCNIRDMAInfiniBandSlurmKubernetes-native Batch SchedulingActionable Alerting SystemsModel Serving InfrastructureInference-load Autoscaling
Industry Keywords
AI/ML WorkloadsHigh-Performance ComputingCNCF ProjectsKubernetes SIGsManaged Kubernetes Services
Tech Stack
Tools & technologiesCloudDistributed SystemsGrafanaKubernetesLinuxPrometheusPythonGo
About the role
Key responsibilities & impact- Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers
- Develop Go/Python automation for end-to-end cluster lifecycle management, including provisioning, upgrades, patching, and deletion
- Build GPU-aware orchestration systems supporting GPU scheduling and resource allocation
- Partner with the Network team on CNI integration, high-performance fabrics, RDMA, and GPUDirect
- Build resilient systems handling timeouts, retries, backoff, and degraded-mode operation across distributed environments
- Develop inference platform services, including model serving infrastructure, inference-load autoscaling, and multi-model deployment patterns
- Build internal tools and CLIs for ML/AI teams to deploy and monitor inference services
- Support and debug production issues through an on-call rotation
- Contribute to Lambda's managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, and higher-level inference and AIOps platform services
- Collaborate with NVIDIA's open-source ecosystem and internal teams across the technology stack
Requirements
What you’ll need- 6+ years of experience in software engineering
- Track record of owning significant technical scope within a team
- Deep understanding of Kubernetes internals, including controllers, schedulers, operators, CRDs, CSI, and CNI
- Solid grasp of distributed systems fundamentals, including fault tolerance, graceful degradation, and failure handling
- Experience operating the control plane and low-level pieces of large-scale Kubernetes clusters
- Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting systems
- Strong programming skills in Go and Python
- Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
- Experience building and operating managed Kubernetes services or Kubernetes control plane components
- Hands-on experience with NVIDIA's GPU/networking ecosystem preferred
- Familiarity with HPC and job schedulers such as Slurm and Kubernetes-native batch scheduling preferred
- Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes preferred
- Exposure to storage architecture for AI/ML workloads preferred
- Past contributions to CNCF projects or Kubernetes SIGs are a plus
- Ability and willingness to work onsite at the San Francisco office 4 days a week
- Legally authorized to work in the United States
Benefits
Comp & perks- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan
- Work from home day currently Tuesday
- Global platform and cross-stack technical exposure
- NVIDIA partnership and access to cutting-edge open-source tooling
- World-class team of engineers with deep expertise in ML, systems, and infrastructure
