FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Expertise in developing optimized compute kernels for deep learning operations, with strong proficiency in C++ programming and performance-oriented parallel programming. Demonstrated ability to collaborate with cross-functional teams to enhance deep learning performance on NVIDIA GPUs.
Highest-signal resume keywords
C++ ProgrammingDeep Learning OptimizationPerformance-Oriented Parallel ProgrammingComputer ArchitectureNVIDIA GPU Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Compute KernelsMatrix MultiplicationConvolutionsNormalizationsRegression TestingCI/CD FlowsPerformance AnalysisDebuggingAssembly ProgrammingResource Optimization
Tools & Technologies
NVIDIA cuDNNNVIDIA cuBLASTensorRTOpenMPPthreads
Industry Keywords
Deep LearningLinear AlgebraPerformance OptimizationThroughput ImprovementSoftware Engineering Best Practices
Tech Stack
Tools & technologiesAssemblyC++
About the role
Key responsibilities & impact- Write highly tuned compute kernels for core deep learning operations, including matrix multiplies, convolutions, and normalizations
- Follow software engineering best practices, including regression testing and CI/CD flows
- Collaborate with the CUDA compiler team to generate optimal assembly code
- Collaborate with deep learning training and inference performance teams to identify layers requiring optimization
- Collaborate with hardware and architecture teams on the programming model for new deep learning hardware features
- Develop optimized code to accelerate linear algebra and deep learning operations on NVIDIA GPUs
- Deliver high-performance code to NVIDIA cuDNN, cuBLAS, and TensorRT libraries
Requirements
What you’ll need- Masters or PhD degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or related field
- 2+ years of relevant industry experience
- Demonstrated strong C++ programming and software design skills, including debugging, performance analysis, and test design
- Experience with performance-oriented parallel programming, even if it’s not on GPUs (e.g. with OpenMP or pthreads)
- Solid understanding of computer architecture and some experience with assembly programming
- Ability to identify bottlenecks, optimize resource utilization, and improve throughput
