FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing ML infrastructure and data platforms, with a strong focus on evaluation standards, reproducibility, and integration across various research efforts. Proficient in Python and PyTorch, with hands-on experience in large-scale GPU workloads and data pipeline design.
Highest-signal resume keywords
ML Infrastructure DevelopmentPython ProgrammingPyTorch FrameworkKubernetes Workload ManagementData Pipeline Design
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
ML InfrastructureData PlatformsPythonPyTorchKubernetesData Pipeline DesignExperimental StatisticsVersioningReproducibilityEvaluation Systems
Soft Skills
Systems ThinkingPragmatic ApproachSelf-StarterHumilityOpen Mindedness
Tools & Technologies
Command Line InterfacesApplication Programming InterfacesGraphical User InterfacesStorage LayersEvaluation Suites
Industry Keywords
Model EvaluationReproducibility StandardsBenchmarking SystemsGenerative ModelsRobotics Policies
Tech Stack
Tools & technologiesKubernetesPythonPyTorch
About the role
Key responsibilities & impact- Own the evaluation platform end to end, including tooling and systems to generate, annotate, review, and adapt evaluations at frontier scale
- Define CLIs, APIs, GUIs, and storage layers to make evaluation seamless, fast, sophisticated, and collaborative
- Work directly with research teams on video, image, audio, agents, and robotics to understand measurement needs and build generalized platform solutions
- Set company-wide standards for model evaluation, including reproducibility, metric definitions, reporting formats, and result trustworthiness
- Support adoption and integration of the platform across training, production model serving, and other research efforts
- Contribute to ML Platform tools and systems that help Runway train and serve frontier models
Requirements
What you’ll need- 5+ years of experience building ML infrastructure or data platforms in production environments
- At least some experience with evaluation, experimentation, or benchmarking systems
- Strong Python and PyTorch skills
- Hands-on experience running large batch GPU workloads on Kubernetes
- Experience designing data pipelines and storage for large volumes of media or model outputs
- Attention to versioning and reproducibility
- Familiarity with experimental statistics, including paired comparisons, confidence intervals, multiple-comparison pitfalls, and inter-rater agreement
- Comfort building internal tools end to end, from the command line to the browser
- Ability to lead a broad technical area, gather requirements, set direction, make tradeoffs, and drive a roadmap
- Familiarity with the full model development lifecycle: data, training, evaluation, and serving
- Self-starter able to work embedded with research teams and move fast
- Strong systems thinking and pragmatic approach to production reliability
- Humility and open mindedness
- Nice to have: experience building evaluation suites for generative models
- Nice to have: hands-on work with LLM- or VLM-as-judge pipelines
- Nice to have: experience with online experimentation platforms and connecting offline metrics to product outcomes
- Nice to have: prior work evaluating agents or robotics policies
Benefits
Comp & perks- Salary range of $240K–$290K for candidates in the U.S.
- Access to best-in-class models, agents, GPUs, storage and cloud services
- Equal opportunity workplace committed to diversity and inclusion
