Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Hewlett Packard Enterprise

Principal Software Engineer – Inference

Hewlett Packard Enterprise

. Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .

Posted 9/16/2026full-timeSpring • Colorado • United StatesLead💰 $152,000 - $349,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in LLM inference runtime development and deployment, with a strong focus on performance optimization, distributed inferencing strategies, and orchestration within enterprise environments. Proficient in mentoring engineering teams and leading architectural reviews to drive technical direction.

Highest-signal resume keywords
LLM Inference Runtime DevelopmentKubernetes Platform ArchitecturesGo ProgrammingPython ProgrammingContinuous Batching

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
LLM Inference EnginesQuantizationTensor ParallelismPipeline ParallelismKV Cache ManagementC++/CUDA DebuggingPerformance EngineeringProfiling Multi-Tier ApplicationsNCCL Collective OperationsSpeculative Decoding
Soft Skills
Analytical AbilitiesProblem-Solving SkillsMentoring
Tools & Technologies
NsightRDMAGPUDirect StorageInfiniBandRoCE
Industry Keywords
Software EngineeringProduction Model ServingEnterprise Software DeliveryAir-Gapped EnvironmentsSovereign Environments

Tech Stack

Tools & technologies
KubernetesPythonC++Go

About the role

Key responsibilities & impact
  • Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference performance engineering teams and own time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, deciding whether to adopt, develop in-house, or decline them
  • Define the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences
  • Lead the model runtime within HPE AI Essentials, an inference platform for enterprises operating large language models on owned infrastructure, including air-gapped and sovereign environments

Requirements

What you’ll need
  • Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving
  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Comprehensive understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect characteristics
  • Expert proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python
  • Ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Experience debugging and profiling multi-tier application workloads such as RAG and Agents
  • Excellent analytical, debugging, and problem-solving abilities
  • Degree in Computer Science or related field
  • Preferred: upstream contributions to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe
  • Preferred: disaggregated prefill/decode serving or KV cache offload and reuse at scale
  • Preferred: RDMA, GPUDirect Storage, InfiniBand, or RoCE
  • Preferred: MIG, fractional GPU allocation, and multi-tenant GPU isolation
  • Preferred: on-premises, air-gapped, or regulated enterprise software delivery

Benefits

Comp & perks
  • Comprehensive health, financial, and emotional wellbeing benefits
  • Personal and professional development programs
  • Flexible work and personal needs management
  • Inclusive workplace culture
  • Employee benefits information available through HPE Rewards