FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Software Engineer – Inference
Hewlett Packard Enterprise. Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in LLM inference runtime development and deployment, with a strong focus on performance optimization, distributed inferencing strategies, and orchestration within enterprise environments. Proficient in mentoring engineering teams and leading architectural reviews to drive technical direction.
Highest-signal resume keywords
LLM Inference Runtime DevelopmentKubernetes Platform ArchitecturesGo ProgrammingPython ProgrammingContinuous Batching
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
LLM Inference EnginesQuantizationTensor ParallelismPipeline ParallelismKV Cache ManagementC++/CUDA DebuggingPerformance EngineeringProfiling Multi-Tier ApplicationsNCCL Collective OperationsSpeculative Decoding
Soft Skills
Analytical AbilitiesProblem-Solving SkillsMentoring
Tools & Technologies
NsightRDMAGPUDirect StorageInfiniBandRoCE
Industry Keywords
Software EngineeringProduction Model ServingEnterprise Software DeliveryAir-Gapped EnvironmentsSovereign Environments
Tech Stack
Tools & technologiesKubernetesPythonC++Go
About the role
Key responsibilities & impact- Define and own the technical direction of LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Partner with inference performance engineering teams and own time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
- Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, deciding whether to adopt, develop in-house, or decline them
- Define the orchestration layer, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences
- Lead the model runtime within HPE AI Essentials, an inference platform for enterprises operating large language models on owned infrastructure, including air-gapped and sovereign environments
Requirements
What you’ll need- Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving
- Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
- Comprehensive understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect characteristics
- Expert proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
- Strong programming proficiency in Go and Python
- Ability to read, debug, and profile C++/CUDA using tools such as Nsight
- Experience debugging and profiling multi-tier application workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving abilities
- Degree in Computer Science or related field
- Preferred: upstream contributions to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe
- Preferred: disaggregated prefill/decode serving or KV cache offload and reuse at scale
- Preferred: RDMA, GPUDirect Storage, InfiniBand, or RoCE
- Preferred: MIG, fractional GPU allocation, and multi-tenant GPU isolation
- Preferred: on-premises, air-gapped, or regulated enterprise software delivery
Benefits
Comp & perks- Comprehensive health, financial, and emotional wellbeing benefits
- Personal and professional development programs
- Flexible work and personal needs management
- Inclusive workplace culture
- Employee benefits information available through HPE Rewards