Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Runway

Member of Technical Staff – ML Platform

Runway

. Own the evaluation platform end to end, including tooling and systems to generate, annotate, review, and adapt evaluations at frontier scale .

Posted 9/22/2026full-timeRemote • United StatesLead💰 $240,000 - $290,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and managing ML infrastructure and data platforms, with a strong focus on evaluation standards, reproducibility, and integration across various research efforts. Proficient in Python and PyTorch, with hands-on experience in large-scale GPU workloads and data pipeline design.

Highest-signal resume keywords
ML Infrastructure DevelopmentPython ProgrammingPyTorch FrameworkKubernetes Workload ManagementData Pipeline Design

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
ML InfrastructureData PlatformsPythonPyTorchKubernetesData Pipeline DesignExperimental StatisticsVersioningReproducibilityEvaluation Systems
Soft Skills
Systems ThinkingPragmatic ApproachSelf-StarterHumilityOpen Mindedness
Tools & Technologies
Command Line InterfacesApplication Programming InterfacesGraphical User InterfacesStorage LayersEvaluation Suites
Industry Keywords
Model EvaluationReproducibility StandardsBenchmarking SystemsGenerative ModelsRobotics Policies

Tech Stack

Tools & technologies
KubernetesPythonPyTorch

About the role

Key responsibilities & impact
  • Own the evaluation platform end to end, including tooling and systems to generate, annotate, review, and adapt evaluations at frontier scale
  • Define CLIs, APIs, GUIs, and storage layers to make evaluation seamless, fast, sophisticated, and collaborative
  • Work directly with research teams on video, image, audio, agents, and robotics to understand measurement needs and build generalized platform solutions
  • Set company-wide standards for model evaluation, including reproducibility, metric definitions, reporting formats, and result trustworthiness
  • Support adoption and integration of the platform across training, production model serving, and other research efforts
  • Contribute to ML Platform tools and systems that help Runway train and serve frontier models

Requirements

What you’ll need
  • 5+ years of experience building ML infrastructure or data platforms in production environments
  • At least some experience with evaluation, experimentation, or benchmarking systems
  • Strong Python and PyTorch skills
  • Hands-on experience running large batch GPU workloads on Kubernetes
  • Experience designing data pipelines and storage for large volumes of media or model outputs
  • Attention to versioning and reproducibility
  • Familiarity with experimental statistics, including paired comparisons, confidence intervals, multiple-comparison pitfalls, and inter-rater agreement
  • Comfort building internal tools end to end, from the command line to the browser
  • Ability to lead a broad technical area, gather requirements, set direction, make tradeoffs, and drive a roadmap
  • Familiarity with the full model development lifecycle: data, training, evaluation, and serving
  • Self-starter able to work embedded with research teams and move fast
  • Strong systems thinking and pragmatic approach to production reliability
  • Humility and open mindedness
  • Nice to have: experience building evaluation suites for generative models
  • Nice to have: hands-on work with LLM- or VLM-as-judge pipelines
  • Nice to have: experience with online experimentation platforms and connecting offline metrics to product outcomes
  • Nice to have: prior work evaluating agents or robotics policies

Benefits

Comp & perks
  • Salary range of $240K–$290K for candidates in the U.S.
  • Access to best-in-class models, agents, GPUs, storage and cloud services
  • Equal opportunity workplace committed to diversity and inclusion