Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Software Engineering Intern, DLFW Comms

NVIDIA

. Integrate new communication library features into AI frameworks from proof of concept through performance analysis to production .

Posted 9/20/2026full-timeShanghai • ChinaEntry LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI frameworks and communication library integration, with a strong focus on kernel authoring and performance optimization on NVIDIA platforms. Proficient in rapid prototyping and development using Python, C++, and CUDA, with a solid understanding of large-scale AI workloads and models.

Highest-signal resume keywords
Kernel AuthoringCUDA OptimizationAI FrameworksPython DevelopmentMulti-GPU Communication

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDAC++PythonTritonCuTePyTorchJAXTRT-LLMVLLMSGLang
Soft Skills
AdaptabilityEffective CommunicationCollaboration
Tools & Technologies
NVIDIA PlatformsAI ModelsPerformance Analysis
Industry Keywords
AI WorkloadsLarge-Scale TrainingProduction InferenceDeep Learning Communication PatternsExpert Parallelism

Tech Stack

Tools & technologies
PythonPyTorchC++

About the role

Key responsibilities & impact
  • Integrate new communication library features into AI frameworks from proof of concept through performance analysis to production
  • Analyze AI workloads and frameworks to identify multi-GPU communication requirements and opportunities
  • Collaborate hands-on with teams working on the latest AI models
  • Author custom communication or fused compute-communication kernels for maximum performance on NVIDIA platforms
  • Conduct in-depth research to achieve SOL GPU performance
  • Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads
  • Collaborate with a dynamic team across multiple time zones

Requirements

What you’ll need
  • Pursuing an M.S. or Ph.D. in CE/CS/EE
  • Strong background in communication, kernel authoring, and/or AI training/inference
  • Rapid prototyping and development with Python, C++, CUDA, or related DSLs such as Triton and cuTe
  • Solid understanding of LLM models and parallelisms
  • Adaptability and passion to learn new areas and tools
  • Flexibility to work and communicate effectively
  • Development experience with PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL is advantageous
  • Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP, and PP is advantageous
  • Experience with CUDA kernel optimization and profiling is advantageous
  • Experience with large-scale training or production inference stack is advantageous

Benefits

Comp & perks
  • Highly competitive salaries
  • Comprehensive benefits package
  • Benefits for employees and their families