FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Inference SW Engineer
Hewlett Packard Enterprise. Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in LLM inference runtimes and production model serving, with strong capabilities in continuous batching, KV cache management, and distributed execution. Proficient in Kubernetes architectures and programming in Go and Python, with a solid understanding of GPU memory hierarchy and performance optimization.
Highest-signal resume keywords
LLM Inference RuntimesContinuous BatchingKubernetes Platform ArchitecturesProgramming Proficiency in Go and PythonTensor and Pipeline Parallelism
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringContinuous BatchingKV Cache ManagementQuantizationSpeculative DecodingTensor and Pipeline ParallelismProgramming in GoProgramming in PythonC++/CUDA DebuggingProfiling Multi-Tier Applications
Soft Skills
Analytical AbilitiesDebugging SkillsProblem-Solving AbilitiesMentoringCode Review
Tools & Technologies
KubernetesNsightVLLMSGLangTensorRT-LLMTGINVIDIA NIMNCCLRDMAInfiniBand
Industry Keywords
LLM ServingGPU Memory HierarchyDisaggregated Prefill/DecodeKV Cache OffloadRegulated Enterprise Software Delivery
Tech Stack
Tools & technologiesKubernetesPythonC++Go
About the role
Key responsibilities & impact- Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload
- Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption
- Contribute to the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Triage and resolve customer issues end-to-end, identify root causes, and improve systems and processes
- Provide code and design reviews, mentor team members, and lead by example on engineering practices
Requirements
What you’ll need- Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving
- Degree in Computer Science or related field
- Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
- Strong understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and GPU memory hierarchy and interconnect characteristics
- Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
- Strong programming proficiency in Go and Python
- Ability to read, debug, and profile C++/CUDA using tools such as Nsight
- Familiarity with debugging/profiling multi-tier application workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving abilities
- Preferred: Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
- Preferred: Disaggregated prefill/decode serving, or KV cache offload and reuse at scale
- Preferred: RDMA, GPUDirect Storage, InfiniBand, or RoCE
- Preferred: MIG, fractional GPU allocation, and multi-tenant GPU isolation
- Preferred: On-premises, air-gapped, or regulated enterprise software delivery
Benefits
Comp & perks- Comprehensive benefits supporting physical, financial, and emotional wellbeing
- Personal and professional development programs
- Flexible work and personal needs management
- Inclusive workplace
- Variable incentives may also be offered