Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

AI Computing Software Development Intern

NVIDIA

. Build and enhance high-performance LLM inference pipelines .

Posted 9/23/2026full-timeShanghai • ChinaEntry LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and optimizing high-performance LLM inference pipelines, with a strong focus on GPU computing, deep learning software performance, and compiler optimization techniques. Proficient in Python and C++ programming, with a solid understanding of CUDA, TensorRT, and performance profiling.

Highest-signal resume keywords
Python ProgrammingC++ ProgrammingCUDA ProgrammingTensorRT OptimizationPerformance Profiling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
LLM Inference PipelinesModel Execution OptimizationGPU Compute KernelsCompiler OptimizationMemory AllocationOperator FusionMulti-GPU Model ServingDeep Learning Software PerformancePerformance AnalysisParallel Programming
Soft Skills
Problem-SolvingCuriosity for AI SystemsPassion for GPU Computing
Tools & Technologies
TensorRTCUDALLVMMLIRNVIDIA GPUs
Industry Keywords
Computer ScienceComputer EngineeringElectrical EngineeringApplied MathematicsAI Systems

Tech Stack

Tools & technologies
PythonPyTorchC++

About the role

Key responsibilities & impact
  • Build and enhance high-performance LLM inference pipelines
  • Analyze and optimize model execution, scalability, and memory usage
  • Collaborate across framework and research teams to deliver efficient multi-GPU model serving
  • Improve TensorRT compiler graph transformations and code generation for NVIDIA GPUs
  • Develop compiler optimization passes
  • Refine operator fusion and memory allocation
  • Collaborate with CUDA and hardware architecture teams
  • Design and tune GPU compute kernels and DSL implementations for GEMM, MoE, Attention, and Convolution
  • Profile, analyze, and improve CUDA kernel performance to maximize GPU efficiency

Requirements

What you’ll need
  • Pursuing an M.S. or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field
  • Excellent problem-solving ability
  • Curiosity for cutting-edge AI systems
  • Passion for GPU computing and deep learning software performance
  • TensorRT LLM track: Strong Python programming and experience with PyTorch; solid understanding of inference and GPU acceleration
  • TensorRT Compiler track: Proficient in C++, with experience in compiler or performance optimization
  • CuTe DSL & CUDA Kernels track: Skilled in C/C++ and CUDA or parallel programming; familiar with LLVM, MLIR and compiler; understanding of computer architecture and performance profiling/analysis/optimization

Benefits

Comp & perks
  • Diverse, supportive work environment
  • Opportunity to work on AI computing platforms driving innovation across industries worldwide