Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Intel Corporation

AI Infrastructure Engineer

Intel Corporation

. Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs .

Posted 10/9/2026full-timeUnited StatesMid-LevelSenior💰 $170,500 - $315,490 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing performance for state-of-the-art LLMs on Intel GPUs, with a strong foundation in GPU computing and high-performance computing. Proficient in modern C++ and Python, with hands-on experience in writing and optimizing custom GPU kernels and contributing to open-source inference engines.

Highest-signal resume keywords
GPU ComputingHigh-Performance ComputingModern C++PythonCustom GPU Kernels

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Performance OptimizationAttention MechanismsQuantizationOperator FusionSystematic ProfilingRoofline AnalysisScale-Out Inference OrchestrationComplex Systems-Level Code ModificationAI SystemsInference Paradigms
Tools & Technologies
TritonSYCLCUDACUTLASSVLLMSGLangPyTorchLlama.cpp
Industry Keywords
Artificial IntelligenceMachine LearningOpen-Source ContributionsGenAI Workload DataCPU/GPU Architecture

Tech Stack

Tools & technologies
Node.jsPythonPyTorchC++

About the role

Key responsibilities & impact
  • Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs
  • Profile, diagnose, and resolve cross-stack performance bottlenecks
  • Design, write, and optimize custom high-performance kernels for attention mechanisms, MoE, quantization, and operator fusions
  • Upstream architectural improvements and hardware backends into open-source repositories such as vLLM, SGLang, and PyTorch
  • Serve as a bridge between hardware teams and the open-source community
  • Apply roofline analysis and systematic profiling to decompose bottlenecks
  • Partner with architecture and compiler teams to shape future GPU roadmaps using real-world GenAI workload data
  • Advance Intel's AI technology through AI infrastructure and performance optimization

Requirements

What you’ll need
  • Bachelor's degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years of experience; Master's degree and 3+ years of experience; or PhD
  • 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC)
  • Proficiency in modern C++ and Python
  • Comfortable reading and modifying complex systems-level code
  • Understanding of CPU/GPU architecture
  • Understanding of modern LLM architectures and inference paradigms, including attention mechanisms, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation
  • Prior open-source contributions to inference engines such as vLLM, SGLang, PyTorch, or llama.cpp
  • Hands-on experience writing and optimizing custom GPU kernels using Triton, SYCL, CUDA/CUTLASS, or other DSLs
  • Experience with scale-out inference orchestration across multi-node topologies
  • Experience leveraging AI coding agents for workflow and benchmark generation

Benefits

Comp & perks
  • Competitive pay
  • Stock bonuses
  • Health benefits
  • Retirement benefits
  • Vacation benefits
  • Hybrid work model allowing employees to split time between working on-site at their assigned Intel site and off-site