FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesNode.jsPython
About the role
Key responsibilities & impact- Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token
- Build tooling that makes performance measurements rigorous and repeatable
- Profile and diagnose performance problems across the serving stack, from scheduling and memory management to kernels and interconnect
- Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments
- Turn findings into production-ready runtimes, configurations, and defaults
- Collaborate with product and infrastructure teams to shape how inference is offered on Runpod
- Monitor the fast-moving inference ecosystem and evaluate what to adopt, build, or contribute back
- Trace serving engine/runtime bottlenecks and implement fixes when configuration tuning is insufficient
- Own LLM serving performance end to end across models, hardware generations, and workloads
Requirements
What you’ll need- 5+ years of professional system engineering experience
- Deep, hands-on experience with vLLM, SGLang, or a comparable serving engine in production or at serious benchmark scale
- Strong software engineering skills in Python
- Comfortable working in large, performance-critical codebases
- Solid understanding of LLM inference performance, including batching, memory, parallelism, and latency-throughput trade-offs
- Experience with inference optimization techniques such as quantization, speculative decoding, or distributed serving
- Rigor in benchmarking and performance analysis
- Comfort with GPU profiling tools
- Ability to explain results clearly in writing and turn them into decisions
- Eligible to work in the United States
- Must not require employment visa sponsorship
- Preferred: experience writing or tuning GPU kernels in CUDA or Triton
- Preferred: contributions to inference or ML systems projects
- Preferred: experience with multi-node GPU systems and high-speed networking
- Preferred: experience at a company where inference cost and latency were core business metrics
Benefits
Comp & perks- Meaningful equity in a fast-growing company; everyone on the team receives stock options
- Generous medical, dental & vision plans
- Flexible PTO
- Remote work-first arrangement
- $1,200 Home Office & Equipment Stipend
- Passionate team on the cutting edge of AI infrastructure, with culture, learning, and ownership at the heart of how the company scales
