FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in optimizing ML systems and inference infrastructure, focusing on performance improvements in cost and latency. Proficient in model serving techniques and experienced with various serving engines and GPU performance optimization.
Highest-signal resume keywords
ML Systems OptimizationModel Serving ExpertisePython ProficiencyGPU Performance EngineeringProduction Experience with Serving Engines
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cost OptimizationLatency ImprovementKV-Cache ManagementContinuous BatchingSpeculative DecodingQuantizationPrefill WorkloadsConcurrency ManagementCUDANCCL
Tools & Technologies
VLLMSGLangTensorRT-LLM
Industry Keywords
Inference InfrastructurePerformance EngineeringMemory BandwidthBatchingHybrid Work Arrangement
Tech Stack
Tools & technologiesPythonRustC++
About the role
Key responsibilities & impact- Own the cost and performance of the inference stack.
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads using real production traffic.
- Tune routing between internal infrastructure and external providers based on cost, capacity, and performance.
- Work with serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
- Build profiling and measurement systems to identify where time, memory, and compute are being spent.
- Collaborate closely with engineers operating the serving fleet.
Requirements
What you’ll need- 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
- Ability to work in person in the Bay Area as part of a hybrid arrangement.
Benefits
Comp & perks- Flexible work: In-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Adaption Passport: Annual travel stipend to explore a country you've never visited.
- Lunch Stipend: Weekly meal allowance for take-out or grocery delivery.
- Well-Being: Comprehensive medical benefits and generous paid time off.
