Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

AI Computing Software Intern, GPU Kernel Libraries

NVIDIA

. Develop high performance operators on NVIDIA GPUs for cuBLAS, TensorRT, cuDNN, cuSparse and cuTensor libraries .

Posted 9/23/2026full-timeShanghai • ChinaEntry LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing high-performance GPU operators using C/C++ and Python, with a strong focus on optimizing deep learning architectures and performance. Proficient in GPU programming with CUDA and familiar with compiler technologies such as LLVM and MLIR.

Highest-signal resume keywords
C/C++ ProgrammingPython DevelopmentCUDA ProgrammingLLVM ExperienceDeep Learning Optimization

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GPU Kernel DevelopmentPerformance AnalysisSoftware DesignAI Technology AdoptionKernel Authoring
Soft Skills
Problem SolvingCommunicationTeamwork
Tools & Technologies
CuBLASTensorRTCuDNNCuSparseCuTensor
Industry Keywords
GPU ComputingDeep Learning ArchitecturePerformance OptimizationComputer Architecture

Tech Stack

Tools & technologies
PythonC++

About the role

Key responsibilities & impact
  • Develop high performance operators on NVIDIA GPUs for cuBLAS, TensorRT, cuDNN, cuSparse and cuTensor libraries
  • Analyze the performance of various GPU kernels on existing/new architecture
  • Identify bottlenecks and propose creative solutions to improve GPU kernel performance
  • Design and develop software for kernel authoring and shipping
  • Adopt cutting-edge AI technologies in GPU kernel or similar development workflow
  • Explore computer architectures for deep learning and work at the intersection of hardware and software
  • Continuously explore and optimize the performance of operators and fusion operators in deep learning networks
  • Design and develop scalable modular infrastructure for optimized operators used in training and inference

Requirements

What you’ll need
  • Pursuing a B.S., M.S., or PhD degree in computer science (or similar)
  • Strong programming skills in C/C++ and Python development
  • Familiar with GPU programming model and CUDA
  • Good understanding about compiler technologies
  • Experience with LLVM and MLIR
  • Excellent problem solving skills
  • Good communication and teamwork
  • Internship role in GPU computing, deep learning architecture, and performance optimization

Benefits

Comp & perks
  • NVIDIA is widely considered to be one of the technology world’s most desirable employers
  • Opportunity to work with forward-thinking and hardworking people
  • Opportunity to work on cutting-edge AI technologies and GPU kernel development