FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Inference Software Engineer
Hewlett Packard Enterprise. Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in LLM inference runtimes and production model serving, with strong capabilities in continuous batching, KV cache management, and distributed execution. Proficient in Kubernetes architectures and advanced programming in Go and Python, with a focus on optimizing performance and resolving complex issues.
Highest-signal resume keywords
LLM Inference RuntimesContinuous BatchingKubernetes Platform ArchitecturesProgramming Proficiency in Go and PythonDebugging and Profiling C++/CUDA
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringLLM Inference EnginesKV Cache ManagementTensor and Pipeline ParallelismQuantizationSpeculative DecodingNCCL Collective OperationsGPU Memory HierarchyC++/CUDA DebuggingAnalytical Problem-Solving
Soft Skills
MentoringCode ReviewCollaboration
Tools & Technologies
KubernetesNsightRDMAGPUDirect StorageInfiniBandRoCE
Industry Keywords
Disaggregated Prefill/Decode ServingMulti-Tenant GPU IsolationOn-Premises Software DeliveryRegulated Enterprise Software Delivery
Tech Stack
Tools & technologiesKubernetesPythonC++Go
About the role
Key responsibilities & impact- Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
- Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption
- Contribute to the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Triage and resolve customer issues end-to-end, identify root causes, and improve systems and processes to prevent recurrence
- Provide code and design reviews, mentor team members, and lead by example on engineering practices
Requirements
What you’ll need- Minimum of 8 years of experience in Software Engineering
- 1–2+ years working directly on LLM inference runtimes or production model serving
- Degree in Computer Science or related field
- Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
- Strong understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Working knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect characteristics
- Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
- Strong programming proficiency in Go and Python
- Ability to read, debug, and profile C++/CUDA using tools such as Nsight
- Familiarity with debugging/profiling multi-tier application workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving abilities
- Preferred: Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
- Preferred: Disaggregated prefill/decode serving, KV cache offload and reuse at scale, RDMA, GPUDirect Storage, InfiniBand, RoCE, MIG, fractional GPU allocation, multi-tenant GPU isolation, and on-premises, air-gapped, or regulated enterprise software delivery
Benefits
Comp & perks- Comprehensive health and wellbeing benefits supporting physical, financial and emotional wellbeing
- Personal and professional development programs
- Flexible work arrangements
- Hybrid work arrangement
- Employee benefits information provided through HPE Rewards