Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Hewlett Packard Enterprise

Senior Inference SW Engineer

Hewlett Packard Enterprise

. Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .

Posted 10/7/2026full-timeUnited StatesSenior💰 $137,000 - $315,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in LLM inference runtimes and production model serving, with strong capabilities in continuous batching, KV cache management, and distributed execution. Proficient in Kubernetes architectures and programming in Go and Python, with a solid understanding of GPU memory hierarchy and performance optimization.

Highest-signal resume keywords
LLM Inference RuntimesContinuous BatchingKubernetes Platform ArchitecturesProgramming Proficiency in Go and PythonTensor and Pipeline Parallelism

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringContinuous BatchingKV Cache ManagementQuantizationSpeculative DecodingTensor and Pipeline ParallelismProgramming in GoProgramming in PythonC++/CUDA DebuggingProfiling Multi-Tier Applications
Soft Skills
Analytical AbilitiesDebugging SkillsProblem-Solving AbilitiesMentoringCode Review
Tools & Technologies
KubernetesNsightVLLMSGLangTensorRT-LLMTGINVIDIA NIMNCCLRDMAInfiniBand
Industry Keywords
LLM ServingGPU Memory HierarchyDisaggregated Prefill/DecodeKV Cache OffloadRegulated Enterprise Software Delivery

Tech Stack

Tools & technologies
KubernetesPythonC++Go

About the role

Key responsibilities & impact
  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption
  • Contribute to the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identify root causes, and improve systems and processes
  • Provide code and design reviews, mentor team members, and lead by example on engineering practices

Requirements

What you’ll need
  • Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Strong understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and GPU memory hierarchy and interconnect characteristics
  • Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python
  • Ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Familiarity with debugging/profiling multi-tier application workloads such as RAG and Agents
  • Excellent analytical, debugging, and problem-solving abilities
  • Preferred: Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
  • Preferred: Disaggregated prefill/decode serving, or KV cache offload and reuse at scale
  • Preferred: RDMA, GPUDirect Storage, InfiniBand, or RoCE
  • Preferred: MIG, fractional GPU allocation, and multi-tenant GPU isolation
  • Preferred: On-premises, air-gapped, or regulated enterprise software delivery

Benefits

Comp & perks
  • Comprehensive benefits supporting physical, financial, and emotional wellbeing
  • Personal and professional development programs
  • Flexible work and personal needs management
  • Inclusive workplace
  • Variable incentives may also be offered