FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Infrastructure Engineer
Intel Corporation. Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in optimizing performance for state-of-the-art LLMs on Intel GPUs, with a strong foundation in GPU computing and high-performance computing. Proficient in modern C++ and Python, with hands-on experience in writing and optimizing custom GPU kernels and contributing to open-source inference engines.
Highest-signal resume keywords
GPU ComputingHigh-Performance ComputingModern C++PythonCustom GPU Kernels
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Performance OptimizationAttention MechanismsQuantizationOperator FusionSystematic ProfilingRoofline AnalysisScale-Out Inference OrchestrationComplex Systems-Level Code ModificationAI SystemsInference Paradigms
Tools & Technologies
TritonSYCLCUDACUTLASSVLLMSGLangPyTorchLlama.cpp
Industry Keywords
Artificial IntelligenceMachine LearningOpen-Source ContributionsGenAI Workload DataCPU/GPU Architecture
Tech Stack
Tools & technologiesNode.jsPythonPyTorchC++
About the role
Key responsibilities & impact- Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs
- Profile, diagnose, and resolve cross-stack performance bottlenecks
- Design, write, and optimize custom high-performance kernels for attention mechanisms, MoE, quantization, and operator fusions
- Upstream architectural improvements and hardware backends into open-source repositories such as vLLM, SGLang, and PyTorch
- Serve as a bridge between hardware teams and the open-source community
- Apply roofline analysis and systematic profiling to decompose bottlenecks
- Partner with architecture and compiler teams to shape future GPU roadmaps using real-world GenAI workload data
- Advance Intel's AI technology through AI infrastructure and performance optimization
Requirements
What you’ll need- Bachelor's degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years of experience; Master's degree and 3+ years of experience; or PhD
- 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC)
- Proficiency in modern C++ and Python
- Comfortable reading and modifying complex systems-level code
- Understanding of CPU/GPU architecture
- Understanding of modern LLM architectures and inference paradigms, including attention mechanisms, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation
- Prior open-source contributions to inference engines such as vLLM, SGLang, PyTorch, or llama.cpp
- Hands-on experience writing and optimizing custom GPU kernels using Triton, SYCL, CUDA/CUTLASS, or other DSLs
- Experience with scale-out inference orchestration across multi-node topologies
- Experience leveraging AI coding agents for workflow and benchmark generation
Benefits
Comp & perks- Competitive pay
- Stock bonuses
- Health benefits
- Retirement benefits
- Vacation benefits
- Hybrid work model allowing employees to split time between working on-site at their assigned Intel site and off-site