FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and evolving Kubernetes-based infrastructures, with a strong focus on GPU scheduling, autoscaling, and CI/CD for machine learning models. Proficient in Python for developing APIs and CLIs, alongside practical experience in AWS and distributed training frameworks.
Highest-signal resume keywords
Kubernetes ExpertisePython DevelopmentAWS Infrastructure ManagementDistributed Training in PyTorchCI/CD Implementation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesPythonAWSPyTorchCI/CDTerraformGPU SchedulingAutoscalingNetworking FundamentalsExperiment Tracking
Soft Skills
Collaborative ApproachUser EmpathyStrong WritingInfluencing Without AuthorityProduct Instincts
Tools & Technologies
HelmArgo WorkflowsRayFlyteKubeflowSlurmBazelNCCLWebDatasetParquet
Industry Keywords
Machine LearningData-Intensive Production SystemsSelf-Driving TechnologiesAutonomous SystemsHigh-Throughput Data Loading
Tech Stack
Tools & technologiesAWSKubernetesNode.jsPythonPyTorchRayTerraform
About the role
Key responsibilities & impact- Build and evolve Kubernetes-based training infrastructure, including GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, operators, and workflow engines
- Shape developer-facing CLIs, SDKs, job submission, templates, and paved paths
- Reduce time to first training run, edit-to-signal latency, and local iteration time; measure and improve these metrics
- Evaluate and evangelize tooling and frameworks through prototypes and migration paths
- Strengthen dataset versioning, sharding, and high-throughput loading for multimodal sensor data
- Turn Python code into tested, documented, observable libraries, CLIs, and services
- Improve experiment hygiene, dashboards, model registry, and dataset-to-result lineage
- Ship CI/CD for models alongside autonomy and simulation
- Build ML-stack observability for utilization, throughput, failures, queue times, and experiment cost
- Develop documentation, onboarding, support, golden-path guides, and office hours
- Drive adoption by prototyping with users, observing workflows, and iterating
- Improve platform reliability, recovery speed, and operational efficiency
- Build self-service security, access-control, data-handling, and cost-governance guardrails with Security, IT, and Infrastructure
Requirements
What you’ll need- 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems
- Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and ability to debug a cluster under load
- Excellent Python and track record of designing APIs and CLIs
- Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar)
- Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling
- Fluency with containers, CI/CD, and modern build systems, including large monorepos
- Ability to influence without authority and persuade senior engineers to adopt changes
- Collaborative approach and user empathy
- Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory
- Passion for self-driving technologies and frontier AI
- Bonus/nice to have: internal developer platform, research platform, or DevEx work
- Bonus/nice to have: large-scale distributed GPU training, NCCL, high-performance cluster networking, and collective communication tuning
- Bonus/nice to have: high-throughput LiDAR or camera data loading, Parquet or WebDataset
- Bonus/nice to have: Argo Workflows, Ray, Flyte, Kubeflow, or Slurm
- Bonus/nice to have: Bazel or similar, including remote caching in a monorepo
- Bonus/nice to have: simulation infrastructure or large-scale batch evaluation pipelines
- Bonus/nice to have: ML, robotics, or autonomous systems infrastructure background
- Bonus/nice to have: security- and IP-sensitive production environments
- Bonus/nice to have: open-source contributions to ML infrastructure or developer tools
Benefits
Comp & perks- Competitive compensation and equity awards
- Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time employees only)
- Unlimited Vacation
- Flexible hours and Work from Home support
- Daily drinks, snacks and catered meals (when in office)
- Regularly scheduled team building activities and social events both on-site, off-site & virtually
- Annual performance bonus
- Workplace accommodations for qualified individuals with disabilities as required by applicable law
