FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Software Engineer – Inference Performance
Baseten. Implement and productionize cutting-edge inference techniques, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling and routing algorithms .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Expertise in implementing and optimizing inference techniques for large language models, with a strong focus on performance improvements and benchmarking across various hardware configurations. Proficient in leveraging ML libraries and programming languages to enhance model architectures and GPU utilization.
Highest-signal resume keywords
Inference OptimizationLLM Optimization TechniquesML Libraries (PyTorch, TensorRT)General-Purpose Programming (Python, C++)GPU Architecture Understanding
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Inference TechniquesQuantizationSpeculative DecodingContinuous BatchingBenchmarking FrameworksModel Architecture TuningPerformance ProfilingRequest SchedulingCache-Aware RoutingLatency/Throughput Optimization
Tools & Technologies
PyTorchTensorRTTensorRT-LLMVLLMSGLang
Industry Keywords
Large Language Models (LLMs)Machine LearningComputer ScienceEngineeringMathematics
Tech Stack
Tools & technologiesPythonPyTorchC++
About the role
Key responsibilities & impact- Implement and productionize cutting-edge inference techniques, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling and routing algorithms
- Profile and optimize inference end to end, from kernel launch overhead and memory layout through request scheduling, prefill/decode disaggregation, and cache-aware routing
- Improve tokens per GPU-hour, utilization, and latency/throughput/cost tradeoffs
- Bring up and tune new model architectures on new hardware
- Build benchmarking frameworks across model architectures, batch sizes, sequence lengths, and hardware configurations
- Contribute upstream to open-source inference engines such as vLLM, SGLang, and TensorRT-LLM
- Partner with model, infrastructure, and customer-facing teams to ship performance improvements
Requirements
What you’ll need- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field
- Experience with one or more general-purpose programming languages, such as Python or C++
- Familiarity with LLM optimization techniques, such as quantization, speculative decoding, and continuous batching
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM
- Demonstrated interest and experience in LLMs
- Deep understanding of GPU architecture
Benefits
Comp & perks- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- (U.S. only) Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities