Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Waabi

Senior/Staff ML Ops Engineer

Waabi

. Build and evolve Kubernetes-based training infrastructure, including GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, operators, and workflow engines .

Posted 9/24/2026full-timeUnited StatesSenior💰 $184,000 - $272,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and evolving Kubernetes-based infrastructures, with a strong focus on GPU scheduling, autoscaling, and CI/CD for machine learning models. Proficient in Python for developing APIs and CLIs, alongside practical experience in AWS and distributed training frameworks.

Highest-signal resume keywords
Kubernetes ExpertisePython DevelopmentAWS Infrastructure ManagementDistributed Training in PyTorchCI/CD Implementation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesPythonAWSPyTorchCI/CDTerraformGPU SchedulingAutoscalingNetworking FundamentalsExperiment Tracking
Soft Skills
Collaborative ApproachUser EmpathyStrong WritingInfluencing Without AuthorityProduct Instincts
Tools & Technologies
HelmArgo WorkflowsRayFlyteKubeflowSlurmBazelNCCLWebDatasetParquet
Industry Keywords
Machine LearningData-Intensive Production SystemsSelf-Driving TechnologiesAutonomous SystemsHigh-Throughput Data Loading

Tech Stack

Tools & technologies
AWSKubernetesNode.jsPythonPyTorchRayTerraform

About the role

Key responsibilities & impact
  • Build and evolve Kubernetes-based training infrastructure, including GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, operators, and workflow engines
  • Shape developer-facing CLIs, SDKs, job submission, templates, and paved paths
  • Reduce time to first training run, edit-to-signal latency, and local iteration time; measure and improve these metrics
  • Evaluate and evangelize tooling and frameworks through prototypes and migration paths
  • Strengthen dataset versioning, sharding, and high-throughput loading for multimodal sensor data
  • Turn Python code into tested, documented, observable libraries, CLIs, and services
  • Improve experiment hygiene, dashboards, model registry, and dataset-to-result lineage
  • Ship CI/CD for models alongside autonomy and simulation
  • Build ML-stack observability for utilization, throughput, failures, queue times, and experiment cost
  • Develop documentation, onboarding, support, golden-path guides, and office hours
  • Drive adoption by prototyping with users, observing workflows, and iterating
  • Improve platform reliability, recovery speed, and operational efficiency
  • Build self-service security, access-control, data-handling, and cost-governance guardrails with Security, IT, and Infrastructure

Requirements

What you’ll need
  • 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems
  • Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and ability to debug a cluster under load
  • Excellent Python and track record of designing APIs and CLIs
  • Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar)
  • Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling
  • Fluency with containers, CI/CD, and modern build systems, including large monorepos
  • Ability to influence without authority and persuade senior engineers to adopt changes
  • Collaborative approach and user empathy
  • Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory
  • Passion for self-driving technologies and frontier AI
  • Bonus/nice to have: internal developer platform, research platform, or DevEx work
  • Bonus/nice to have: large-scale distributed GPU training, NCCL, high-performance cluster networking, and collective communication tuning
  • Bonus/nice to have: high-throughput LiDAR or camera data loading, Parquet or WebDataset
  • Bonus/nice to have: Argo Workflows, Ray, Flyte, Kubeflow, or Slurm
  • Bonus/nice to have: Bazel or similar, including remote caching in a monorepo
  • Bonus/nice to have: simulation infrastructure or large-scale batch evaluation pipelines
  • Bonus/nice to have: ML, robotics, or autonomous systems infrastructure background
  • Bonus/nice to have: security- and IP-sensitive production environments
  • Bonus/nice to have: open-source contributions to ML infrastructure or developer tools

Benefits

Comp & perks
  • Competitive compensation and equity awards
  • Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time employees only)
  • Unlimited Vacation
  • Flexible hours and Work from Home support
  • Daily drinks, snacks and catered meals (when in office)
  • Regularly scheduled team building activities and social events both on-site, off-site & virtually
  • Annual performance bonus
  • Workplace accommodations for qualified individuals with disabilities as required by applicable law