Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Baseten

Engineering Manager – Inference Performance

Baseten

. Lead, mentor and grow a team of inference performance engineers through regular 1:1s, clear feedback, career development and performance reviews .

Posted 10/7/2026full-timeUnited StatesMid-LevelSenior💰 $240,000 - $270,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in GPU optimization, performance engineering, and team leadership, with a strong focus on driving technical roadmaps and delivering measurable outcomes. Proficient in managing and mentoring engineering teams while ensuring high standards of engineering quality and operational excellence.

Highest-signal resume keywords
GPU OptimizationPerformance EngineeringTeam LeadershipTechnical Roadmap ManagementML Libraries Familiarity

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GPU WorkloadsQuantizationSpeculative DecodingCUDATritonCUTLASSPerformance ReviewsInference OptimizationBenchmarkingOperational Excellence
Soft Skills
Clear CommunicationMentoringFeedback DeliveryStakeholder Alignment
Tools & Technologies
PyTorchTensorRTTensorRT-LLMVLLMSGLangBaseten Inference Stack
Industry Keywords
Inference Performance EngineeringModel ArchitecturesTeam ScalingStartup EnvironmentEngineering Quality Standards

Tech Stack

Tools & technologies
PyTorch

About the role

Key responsibilities & impact
  • Lead, mentor and grow a team of inference performance engineers through regular 1:1s, clear feedback, career development and performance reviews
  • Hire GPU and inference engineering talent and build a strong collaborative team culture
  • Own the technical roadmap and execution for runtime performance work
  • Review designs, guide profiling and optimization efforts, and help the team reason about time and memory usage
  • Drive productionization of quantization, speculative decoding, KV-cache reuse, chunked prefill and custom scheduling
  • Turn performance improvements into measurable outcomes including tokens per GPU-hour, utilization, latency and cost
  • Bring up and tune new model architectures on new hardware
  • Partner with Infrastructure, Inference Platform, Kernels, Model APIs and customer-facing teams to set priorities, coordinate launches and ship improvements
  • Set standards for engineering quality, benchmarking, operational excellence and incident response
  • Work on agentic inference optimization, agentic kernels, speculative decoding model training and the Baseten Inference Stack

Requirements

What you’ll need
  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field
  • Experience managing engineers, including hiring, mentoring, giving feedback and running performance reviews
  • Experience leading or closely supporting GPU optimization teams in training, inference or recommendation systems
  • Strong technical depth in GPU workloads, with a solid understanding of GPU architecture and performance tradeoffs
  • Familiarity with ML libraries such as PyTorch, TensorRT or TensorRT-LLM
  • A track record of driving roadmaps and shipping complex technical projects with a team
  • Clear written and verbal communication, including the ability to align stakeholders across teams
  • Familiarity with inference engines such as vLLM, SGLang or TensorRT-LLM
  • Experience with LLM optimization techniques such as quantization, speculative decoding and continuous batching
  • Experience with GPU kernels such as CUDA, Triton and CUTLASS
  • Experience scaling a team through rapid growth at a startup
  • A background as a hands-on performance or systems engineer before moving into management

Benefits

Comp & perks
  • Competitive compensation, including meaningful equity
  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • (U.S. only) Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities