Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Baseten

Software Engineer – Inference Performance

Baseten

. Implement and productionize cutting-edge inference techniques, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling and routing algorithms .

Posted 10/6/2026full-timeUnited StatesMid-LevelSenior💰 $180,000 - $360,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Expertise in implementing and optimizing inference techniques for large language models, with a strong focus on performance improvements and benchmarking across various hardware configurations. Proficient in leveraging ML libraries and programming languages to enhance model architectures and GPU utilization.

Highest-signal resume keywords
Inference OptimizationLLM Optimization TechniquesML Libraries (PyTorch, TensorRT)General-Purpose Programming (Python, C++)GPU Architecture Understanding

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Inference TechniquesQuantizationSpeculative DecodingContinuous BatchingBenchmarking FrameworksModel Architecture TuningPerformance ProfilingRequest SchedulingCache-Aware RoutingLatency/Throughput Optimization
Tools & Technologies
PyTorchTensorRTTensorRT-LLMVLLMSGLang
Industry Keywords
Large Language Models (LLMs)Machine LearningComputer ScienceEngineeringMathematics

Tech Stack

Tools & technologies
PythonPyTorchC++

About the role

Key responsibilities & impact
  • Implement and productionize cutting-edge inference techniques, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling and routing algorithms
  • Profile and optimize inference end to end, from kernel launch overhead and memory layout through request scheduling, prefill/decode disaggregation, and cache-aware routing
  • Improve tokens per GPU-hour, utilization, and latency/throughput/cost tradeoffs
  • Bring up and tune new model architectures on new hardware
  • Build benchmarking frameworks across model architectures, batch sizes, sequence lengths, and hardware configurations
  • Contribute upstream to open-source inference engines such as vLLM, SGLang, and TensorRT-LLM
  • Partner with model, infrastructure, and customer-facing teams to ship performance improvements

Requirements

What you’ll need
  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field
  • Experience with one or more general-purpose programming languages, such as Python or C++
  • Familiarity with LLM optimization techniques, such as quantization, speculative decoding, and continuous batching
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM
  • Demonstrated interest and experience in LLMs
  • Deep understanding of GPU architecture

Benefits

Comp & perks
  • Competitive compensation, including meaningful equity
  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • (U.S. only) Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities