Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Hewlett Packard Enterprise

Senior Software Engineer, Inference

Hewlett Packard Enterprise

. Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .

Posted 9/16/2026full-timeSpring • Colorado • United StatesSenior💰 $137,000 - $315,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in LLM inference runtimes and production model serving, with strong capabilities in continuous batching, KV cache management, and distributed execution. Proficient in Kubernetes architectures and programming in Go and Python, with a solid understanding of GPU memory hierarchies and performance optimization.

Highest-signal resume keywords
LLM Inference RuntimesContinuous BatchingKubernetes Platform ArchitecturesProgramming Proficiency in Go and PythonDebugging and Profiling C++/CUDA

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringLLM Inference EnginesKV Cache ManagementTensor and Pipeline ParallelismQuantizationSpeculative DecodingNCCL Collective OperationsDebugging Multi-Tier ApplicationsRDMAGPUDirect Storage
Soft Skills
Analytical AbilitiesProblem-SolvingMentoring
Tools & Technologies
KubernetesNsightVLLMSGLangTensorRT-LLMTGINVIDIA NIMLMCacheKServeInfiniBand
Industry Keywords
Disaggregated Prefill/Decode ServingMulti-Tenant GPU IsolationOn-Premises Software DeliveryRegulated Enterprise Software Delivery

Tech Stack

Tools & technologies
KubernetesPythonC++Go

About the role

Key responsibilities & impact
  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption
  • Contribute to the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identify root causes, and improve systems and processes
  • Provide code and design reviews, mentor team members, and lead by example on engineering practices

Requirements

What you’ll need
  • Minimum of 8 years of experience in Software Engineering
  • 1-2+ years working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Strong understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and GPU memory hierarchy and interconnect characteristics
  • Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python
  • Ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Familiarity with debugging/profiling multi-tier application workloads such as RAG and Agents
  • Excellent analytical, debugging, and problem-solving abilities
  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe preferred
  • Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale preferred
  • Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE preferred
  • Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation preferred
  • Experience with on-premises, air-gapped, or regulated enterprise software delivery preferred

Benefits

Comp & perks
  • Comprehensive health and wellbeing benefits supporting physical, financial, and emotional wellbeing
  • Personal and professional development programs
  • Career development opportunities
  • Flexible work arrangements to manage work and personal needs
  • Inclusive workplace culture
  • Reasonable accommodation during the application or interview process for qualified applicants with disabilities