FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Software Engineer, Inference
Hewlett Packard Enterprise. Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in LLM inference runtimes and production model serving, with a strong focus on performance optimization, distributed inferencing strategies, and technical leadership. Proficient in Kubernetes architectures and advanced programming in Go and Python, with a solid foundation in debugging and profiling complex applications.
Highest-signal resume keywords
LLM Inference RuntimesKubernetes Platform ArchitecturesGo ProgrammingPython ProgrammingContinuous Batching
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringLLM Inference EnginesQuantizationTensor ParallelismPipeline ParallelismKV Cache ManagementC++/CUDA DebuggingPerformance EngineeringProfiling Multi-Tier ApplicationsNCCL Collective Operations
Soft Skills
Analytical AbilitiesProblem-SolvingMentoringTechnical Presentation
Tools & Technologies
NsightRDMAGPUDirect StorageInfiniBandRoCE
Industry Keywords
Large Language ModelsInference PerformanceDisaggregated Prefill/DecodeAutoscalingAir-Gapped Environments
Tech Stack
Tools & technologiesKubernetesPythonC++Go
About the role
Key responsibilities & impact- Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Partner with inference performance engineering teams and own time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
- Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, determining whether to adopt, develop in-house, or decline them
- Define the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences
- Lead the model runtime for HPE AI Essentials, an inference platform enabling enterprises to operate large language models on owned infrastructure, including air-gapped and sovereign environments
Requirements
What you’ll need- Minimum of 12 years of experience in Software Engineering
- 1+ years working directly on LLM inference runtimes or production model serving
- Degree in Computer Science or related field
- Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
- Comprehensive understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect characteristics
- Expert proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
- Strong programming proficiency in Go and Python
- Ability to read, debug, and profile C++/CUDA using tools such as Nsight
- Experience debugging and profiling multi-tier application workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving abilities
- Preferred: upstream contributions to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe
- Preferred: disaggregated prefill/decode serving or KV cache offload and reuse at scale
- Preferred: RDMA, GPUDirect Storage, InfiniBand, or RoCE
- Preferred: MIG, fractional GPU allocation, and multi-tenant GPU isolation
- Preferred: on-premises, air-gapped, or regulated enterprise software delivery
Benefits
Comp & perks- Comprehensive health and wellbeing benefits supporting physical, financial, and emotional wellbeing
- Personal and professional development programs
- Flexible work arrangements to manage work and personal needs
- Inclusive workplace culture
- Reasonable accommodation during the application or interview process for applicants with disabilities