Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Performance Software Intern, Deep Learning Libraries

NVIDIA

. Write highly tuned compute kernels to perform core deep learning operations, including matrix multiplies, MoE, and Attention .

Posted 9/23/2026full-timeShanghai • ChinaEntry LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing deep learning operations and performance tuning on NVIDIA GPUs, with a strong foundation in parallel programming and computer architecture. Proficient in software engineering best practices, including regression testing and CI/CD workflows.

Highest-signal resume keywords
CUDA GPU ProgrammingPerformance-Oriented Parallel ProgrammingDeep Learning Library Kernel TuningComputer Architecture UnderstandingAssembly Programming Experience

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Deep Learning OperationsMatrix MultiplicationPerformance AnalysisDebuggingNumerical MethodsLinear AlgebraTuning AlgorithmsOptimized Code DevelopmentRegression TestingTest Design
Tools & Technologies
OpenMPPthreadsLLVMTVM Tensor ExpressionsTensorFlow MLIR
Industry Keywords
Deep LearningCompute KernelsCI/CDAssembly Code GenerationHardware Programming Model

Tech Stack

Tools & technologies
AssemblyTensorflow

About the role

Key responsibilities & impact
  • Write highly tuned compute kernels to perform core deep learning operations, including matrix multiplies, MoE, and Attention
  • Follow software engineering best practices, including regression testing and CI/CD flows
  • Collaborate with the Compiler team on generating optimal assembly code
  • Collaborate with deep learning training and inference performance teams to identify layers requiring optimization
  • Collaborate with hardware and architecture teams on the programming model for new deep learning hardware features
  • Develop optimized code to accelerate linear algebra and deep learning operations on NVIDIA GPUs
  • Tune parallel algorithms and analyze their performance

Requirements

What you’ll need
  • Pursuing Masters or PhD degree in Computer Science, Computer Engineering, Applied Math, or related field
  • Demonstrated strong programming and software design skills, including debugging, performance analysis, and test design
  • Experience with performance-oriented parallel programming, even if it’s not on GPUs (e.g. with OpenMP or pthreads)
  • Solid understanding of computer architecture and some experience with assembly programming
  • Ability to identify bottlenecks, optimize resource utilization, and improve throughput
  • Tuning deep learning library kernel code
  • CUDA GPU programming
  • Numerical methods and linear algebra
  • LLVM, TVM tensor expressions, or TensorFlow MLIR