Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
SPREEAI

MLOps Engineer

SPREEAI

. Design and operate training-as-a-service infrastructure .

Posted 9/16/2026full-timeSan Francisco • California • United StatesMid-LevelSenior💰 $145,000 - $180,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating machine learning platforms, with a strong focus on training orchestration, experiment tracking, and data versioning. Proficient in Python or Go, and experienced in managing multi-GPU training jobs and infrastructure in a fast-paced environment.

Highest-signal resume keywords
ML Platform DevelopmentTraining Orchestration Using RayDocker ExperienceKubernetes ManagementExperiment Tracking and Data Versioning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonGoMulti-GPU TrainingCI/CDExperiment TrackingData VersioningTraining Job MonitoringCost GovernanceA/B TestingInfrastructure Tooling
Soft Skills
Mentoring EngineersOwnership of Platform Direction
Tools & Technologies
RayKubeflowAirflowDockerKubernetesHelmByteplusFireworks.AI
Industry Keywords
Training-as-a-ServiceML LifecycleSpot InstancesPreemptible VMsBudget Alerts

Tech Stack

Tools & technologies
AirflowDockerKubernetesPythonRayGo

About the role

Key responsibilities & impact
  • Design and operate training-as-a-service infrastructure
  • Enable scientists to launch multi-GPU training jobs, track metrics, and receive completion notifications without directly managing infrastructure
  • Build model CI/CD with automated evaluation gates, canary rollouts, and A/B testing hooks
  • Own experiment tracking and data versioning for terabyte-scale datasets that change weekly
  • Monitor training job health, including GPU utilization, loss curves, and OOM detection
  • Drive training cost governance through spot instances, preemptible VMs, and budget alerts
  • Partner with ML Scientists to translate workflow pain points into platform abstractions
  • Evaluate and integrate external model providers, including Byteplus and Fireworks.AI, into the training and evaluation platform

Requirements

What you’ll need
  • Experience building or operating an ML platform
  • Experience with training orchestration using Ray, Kubeflow, Airflow, or a custom solution
  • Experience with Docker
  • Experience with Kubernetes, including Jobs/CronJobs and Helm
  • Proficiency in Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction and mentoring engineers
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment

Benefits

Comp & perks
  • 🌐 Worldwide ❌ Jobs You've Hidden ⭐️ Saved Jobs ✅ Applied Jobs ✉️ Email Alerts 👤 Account SPREEAI Website LinkedIn All Job Openings 11 - 50 employees 👗 Fashion 🤖 Artificial Intelligence 🛍️ eCommerce 💰 Funding Round - SpreeAI on 2025-05 Fashion
  • Artificial Intelligence
  • eCommerce SPREEAI is an AI-driven fashion technology company that provides real-time virtual try-on experiences for apparel, embedded directly into a brand's shopping platform. Customers can snap or upload a photo or select a model to see garments styled and fitted instantly, helping brands increase personalization and conversion. SPREEAI partners with fashion labels, industry organizations, and academic institutions to advance its fit and visualization technology within e-commerce. MLOps Engineer Job not on LinkedIn 🔥 1 hour ago 🏢🏡 San Francisco – Hybrid 💵 $145k - $180k / year ⏰ Full Time 🟡 Mid-level 🟠 Senior 🤖 Machine Learning Engineer 👻 Ghost score 10% Airflow Docker Kubernetes Python Ray Go Apply Now Customize resume + cover letter Report problem ☆ Save ☑️ Mark as applied ❌ Hide 📋 Description
  • Design and operate training-as-a-service infrastructure
  • Enable scientists to launch multi-GPU training jobs, track metrics, and receive completion notifications without directly managing infrastructure
  • Build model CI/CD with automated evaluation gates, canary rollouts, and A/B testing hooks
  • Own experiment tracking and data versioning for terabyte-scale datasets that change weekly
  • Monitor training job health, including GPU utilization, loss curves, and OOM detection
  • Drive training cost governance through spot instances, preemptible VMs, and budget alerts
  • Partner with ML Scientists to translate workflow pain points into platform abstractions
  • Evaluate and integrate external model providers, including Byteplus and Fireworks.AI, into the training and evaluation platform 🎯 Requirements
  • Experience building or operating an ML platform
  • Experience with training orchestration using Ray, Kubeflow, Airflow, or a custom solution
  • Experience with Docker
  • Experience with Kubernetes, including Jobs/CronJobs and Helm
  • Proficiency in Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction and mentoring engineers
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment Apply Now 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score Similar Jobs Senior Machine Learning Engineer, Fraud 🕒 3 days ago Plaid 501 - 1000 🔌 API 💳 Fintech 🤝 B2B Website LinkedIn All Job Openings Senior Machine Learning Engineer building production ML models and pipelines for Plaid’s fraud detection products. Improving financial security across Plaid’s network of institutions and customers. 🏢🏡 San Francisco – Hybrid 💵 $229k - $315.4k / year ⏰ Full Time 🟠 Senior 🤖 Machine Learning Engineer Python PyTorch Scikit-Learn SQL Senior AI/ML Engineer 🕒 September 4 Metriport 11 - 50 🏥 Healthcare 🔌 API ☁️ SaaS Website LinkedIn All Job Openings Senior AI/ML Engineer building predictive models and ML infrastructure for Metriport’s medical data exchange platform. Turning messy clinical records into actionable intelligence for healthcare customers. 🏢🏡 San Francisco – Hybrid 💵 $200k - $260k / year 💰 $2.4M Seed Round - Metriport on 2022-12 ⏰ Full Time 🟠 Senior 🤖 Machine Learning Engineer AWS Cloud SQL Machine Learning Research Manager 🕒 September 3 Rad AI 51 - 200 🏥 Healthcare 🏭 Manufacturing 🔧 Hardware Website LinkedIn All Job Openings Machine Learning Research Manager leading applied and clinical research at Rad AI, transforming radiology with artificial intelligence. Guiding NLP, LLM, and clinical AI systems from research to production. 🏢🏡 San Francisco – Hybrid 💰 $25M Series A on 2021-11 ⏰ Full Time 🟡 Mid-level 🟠 Senior 🤖 Machine Learning Engineer 🦅 H1B Visa Sponsor Cloud PyTorch Machine Learning Engineer II, Responsible AI 🕒 August 27 Pinterest 1001 - 5000 📱 Media 👥 B2C Website LinkedIn All Job Openings Machine Learning Engineer II advancing responsible AI, fairness, and generative AI safeguards at Pinterest. Collaborating across engineering to improve its visual discovery platform. 🏢🏡 San Francisco – Hybrid 💵 $138.9k - $286k / year 💰 Post IPO equity on 2022-08 ⏰ Full Time 🟢 Junior 🟡 Mid-level 🤖 Machine Learning Engineer 🦅 H1B Visa Sponsor Software Engineer, ML Developer 🕒 August 27 Anyscale 51 - 200 🤖 Artificial Intelligence ☁️ SaaS 🏢 Enterprise Website LinkedIn All Job Openings Software Engineer building Ray developer tools, ML workflows, and observability. Operating scalable backend services for Anyscale’s distributed computing platform. 🏢🏡 San Francisco – Hybrid 💵 $226k - $283k / year ⏰ Full Time 🟡 Mid-level 🟠 Senior 🤖 Machine Learning Engineer 🦅 H1B Visa Sponsor Cloud Distributed Systems PyTorch Ray View More Machine Learning Engineer Jobs 🌐 Worldwide Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or support@remoterocketship.com Search Remote jobs Search Jobs by country Search jobs by city Search jobs by job title Search entry-level jobs Search junior-level jobs Search senior-level jobs Search jobs by tech stack Search jobs by contract type Search remote internships Search remote part-time jobs Remote jobs Anywhere in the World Companies Hiring Anywhere in the World Companies Hiring Sales People Anywhere in the World Companies Hiring Software Engineers Anywhere in the World Resources About us Advice Tips for finding remote jobs Interview questions and answers Resume examples Cover letter examples Post a job Affiliates Is Remote Rocketship legit? Privacy policy Terms of service Job board SEO course Remote Job Search MasterClass Resume Review AI Apply Copilot OpenClaw job finder API docs Find jobs using your resume Jobs by Country Remote jobs anywhere in the world (Worldwide remote jobs) Remote jobs United States Remote jobs Australia Remote jobs Brazil Remote jobs Canada Remote jobs France Remote jobs Ireland Remote jobs Germany Remote jobs Netherlands Remote jobs Spain Remote jobs UK Popular Jobs Remote data analyst jobs Remote customer support jobs Remote executive assistant jobs Remote marketing jobs Remote product designer jobs Remote product manager jobs Remote project manager jobs Remote recruiter jobs Remote sales jobs Remote software engineer jobs Jobs by Type Remote full-time jobs Remote part-time jobs Remote contract jobs Remote internship jobs Remote entry-level jobs Remote jobs with no experience required Remote junior jobs (1-3 years of experience) Digital nomad jobs Remote jobs with no degree required Freelance remote jobs Temporary remote jobs Remote jobs hiring now Stay at home mom jobs