Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
DigitalOcean

Principal Engineer, Model Optimizations

DigitalOcean

. Set the technical strategy for model optimization across DigitalOcean's heterogeneous NVIDIA and AMD accelerator fleet .

Posted 10/8/2026full-timeSeattle • Washington • United StatesLead💰 $249,600 - $312,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in optimizing LLM inference performance across diverse GPU architectures, with a strong focus on quantization strategies and kernel-level programming. Proven ability to lead technical direction, mentor engineers, and collaborate with cross-functional teams to drive impactful results.

Highest-signal resume keywords
LLM Inference OptimizationCUDA ProgrammingQuantization ExpertiseGPU Architecture UnderstandingCross-Functional Leadership

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDAC++PythonKernel-Level ProgrammingQuantizationPerformance OptimizationBenchmarkingProfilingModel EvaluationTraffic Distribution Analysis
Soft Skills
Excellent CommunicationMentoring
Tools & Technologies
NsightRocprofTritonCUTLASSHIPVLLMSGLangTensorRT-LLMAITERComposable Kernel
Industry Keywords
Performance-Critical SystemsInference PerformanceGPU FamiliesModel OptimizationTechnical Strategy

Tech Stack

Tools & technologies
PythonC++

About the role

Key responsibilities & impact
  • Set the technical strategy for model optimization across DigitalOcean's heterogeneous NVIDIA and AMD accelerator fleet
  • Own quantization strategy, calibration methodology, accuracy budgets, and evaluation gates
  • Drive execution-level performance for MoE, MLA, GQA, long-context, and sliding-window attention architectures
  • Lead speculative decoding work and acceptance-rate tuning
  • Write and tune CUDA, Triton, CUTLASS, HIP, Composable Kernel, hipBLASLt, and AITER kernels where upstream falls short
  • Establish repeatable parallelism layouts across models, GPU families, and traffic shapes
  • Build benchmarking and regression infrastructure for traffic distributions, TTFT, ITL, throughput, and accuracy validation
  • Make AMD a first-class target and drive upstream work in vLLM, SGLang, and TensorRT-LLM
  • Partner with NVIDIA and AMD engineering teams on pre-silicon enablement, early-access hardware, and roadmap feedback
  • Set technical direction, mentor senior and staff engineers, and represent DigitalOcean in upstream communities, conferences, and customer technical deep dives

Requirements

What you’ll need
  • 12+ years in performance-critical systems, with substantial recent experience optimizing LLM inference in production
  • Deep understanding of GPU architecture and inference performance: memory bandwidth versus compute bounds, arithmetic intensity, kernel launch and scheduling overhead, and prefill versus decode differences
  • Hands-on kernel-level experience on at least one vendor stack and ability to work across CUDA/CUTLASS/Triton and ROCm/HIP/CK
  • Practical quantization expertise
  • Familiarity with vLLM, SGLang, or TensorRT-LLM internals sufficient to land non-trivial upstream changes
  • Strong Python and C++/CUDA skills
  • Profiling experience with Nsight, rocprof, or equivalent tooling
  • Measurement-first approach using reproducible benchmarks
  • Excellent written and verbal communication
  • Experience leading cross-functional efforts spanning infrastructure, product, and customers

Benefits

Comp & perks
  • Reimbursement for relevant conferences, training, and education
  • Access to LinkedIn Learning's 10,000+ courses
  • Employee Assistance Program
  • Local Employee Meetups
  • Flexible time off policy
  • Bonus eligibility based on company and individual performance
  • Equity compensation, including equity grants upon hire
  • Employee Stock Purchase Program
  • Hybrid work arrangement