Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Hewlett Packard Enterprise

Senior Software Engineer, Inference

Hewlett Packard Enterprise

. Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .

Posted 9/16/2026full-timeSpring • Colorado • United StatesSenior💰 $137,000 - $315,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in LLM inference runtimes and production model serving, with strong capabilities in continuous batching, KV cache management, and distributed execution. Proficient in Kubernetes architectures and advanced programming in Go and Python, with a focus on optimizing performance and resolving complex engineering challenges.

Highest-signal resume keywords
LLM Inference RuntimesContinuous BatchingKubernetes ArchitectureProgramming Proficiency in Go and PythonTensor and Pipeline Parallelism

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringKV Cache ManagementQuantizationSpeculative DecodingNCCL Collective OperationsC++/CUDA DebuggingDisaggregated Prefill/Decode ServingGPU Memory HierarchyProfiling Multi-Tier ApplicationsEngine Internals Modification
Soft Skills
MentoringProblem SolvingCollaboration
Tools & Technologies
KubernetesNsightRDMAGPUDirect StorageInfiniBandRoCEMIGVLLMTensorRT-LLMSGLang
Industry Keywords
Model ServingDistributed ExecutionAutoscalingCustomer Issue ResolutionEngineering Practices

Tech Stack

Tools & technologies
KubernetesPythonC++Go

About the role

Key responsibilities & impact
  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption
  • Contribute to the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes
  • Provide code and design reviews, mentor team members, and lead by example on engineering practices

Requirements

What you’ll need
  • Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Strong understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and GPU memory hierarchy and interconnect characteristics
  • Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python
  • Ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Familiarity with debugging/profiling multi-tier application workloads such as RAG and Agents
  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe preferred
  • Experience with disaggregated prefill/decode serving, KV cache offload and reuse at scale, RDMA, GPUDirect Storage, InfiniBand, RoCE, MIG, fractional GPU allocation, multi-tenant GPU isolation, and on-premises, air-gapped, or regulated enterprise software delivery preferred

Benefits

Comp & perks
  • Comprehensive health and wellbeing benefits supporting physical, financial and emotional wellbeing
  • Personal and professional development programs
  • Flexible work and personal needs management
  • Inclusive workplace
  • Reasonable accommodation during the application or interview process for applicants with disabilities