Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
adaption

Inference Engineer

adaption

. Own the cost and performance of the inference stack.

Posted 10/6/2026full-timeSan Francisco • California • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing ML systems and inference infrastructure, focusing on performance improvements in cost and latency. Proficient in model serving techniques and experienced with various serving engines and GPU performance optimization.

Highest-signal resume keywords
ML Systems OptimizationModel Serving ExpertisePython ProficiencyGPU Performance EngineeringProduction Experience with Serving Engines

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cost OptimizationLatency ImprovementKV-Cache ManagementContinuous BatchingSpeculative DecodingQuantizationPrefill WorkloadsConcurrency ManagementCUDANCCL
Tools & Technologies
VLLMSGLangTensorRT-LLM
Industry Keywords
Inference InfrastructurePerformance EngineeringMemory BandwidthBatchingHybrid Work Arrangement

Tech Stack

Tools & technologies
PythonRustC++

About the role

Key responsibilities & impact
  • Own the cost and performance of the inference stack.
  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads using real production traffic.
  • Tune routing between internal infrastructure and external providers based on cost, capacity, and performance.
  • Work with serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems to identify where time, memory, and compute are being spent.
  • Collaborate closely with engineers operating the serving fleet.

Requirements

What you’ll need
  • 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
  • Ability to work in person in the Bay Area as part of a hybrid arrangement.

Benefits

Comp & perks
  • Flexible work: In-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
  • Adaption Passport: Annual travel stipend to explore a country you've never visited.
  • Lunch Stipend: Weekly meal allowance for take-out or grocery delivery.
  • Well-Being: Comprehensive medical benefits and generous paid time off.