Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Runpod

Senior ML Systems Engineer, Inference

Runpod

. Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token .

Posted 9/25/2026full-timeRemote • United StatesSenior💰 $150,000 - $220,000 per yearWebsite

Tech Stack

Tools & technologies
Node.jsPython

About the role

Key responsibilities & impact
  • Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token
  • Build tooling that makes performance measurements rigorous and repeatable
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management to kernels and interconnect
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments
  • Turn findings into production-ready runtimes, configurations, and defaults
  • Collaborate with product and infrastructure teams to shape how inference is offered on Runpod
  • Monitor the fast-moving inference ecosystem and evaluate what to adopt, build, or contribute back
  • Trace serving engine/runtime bottlenecks and implement fixes when configuration tuning is insufficient
  • Own LLM serving performance end to end across models, hardware generations, and workloads

Requirements

What you’ll need
  • 5+ years of professional system engineering experience
  • Deep, hands-on experience with vLLM, SGLang, or a comparable serving engine in production or at serious benchmark scale
  • Strong software engineering skills in Python
  • Comfortable working in large, performance-critical codebases
  • Solid understanding of LLM inference performance, including batching, memory, parallelism, and latency-throughput trade-offs
  • Experience with inference optimization techniques such as quantization, speculative decoding, or distributed serving
  • Rigor in benchmarking and performance analysis
  • Comfort with GPU profiling tools
  • Ability to explain results clearly in writing and turn them into decisions
  • Eligible to work in the United States
  • Must not require employment visa sponsorship
  • Preferred: experience writing or tuning GPU kernels in CUDA or Triton
  • Preferred: contributions to inference or ML systems projects
  • Preferred: experience with multi-node GPU systems and high-speed networking
  • Preferred: experience at a company where inference cost and latency were core business metrics

Benefits

Comp & perks
  • Meaningful equity in a fast-growing company; everyone on the team receives stock options
  • Generous medical, dental & vision plans
  • Flexible PTO
  • Remote work-first arrangement
  • $1,200 Home Office & Equipment Stipend
  • Passionate team on the cutting edge of AI infrastructure, with culture, learning, and ownership at the heart of how the company scales