FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Deep Learning Software Engineer, Inference, Model Optimization
NVIDIA. Train, develop, and deploy state-of-the-art generative AI models such as LLMs and diffusion models using NVIDIA's AI software stack .
Posted 10/9/2026full-timeRemote • California • United StatesSenior💰 $184,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and optimizing generative AI models, particularly using NVIDIA's AI software stack and the Torch ecosystem. Proficient in high-performance GPU kernel development and performance analysis, with a strong foundation in deep learning and machine learning frameworks.
Highest-signal resume keywords
Generative AI Model DevelopmentDeep Learning ExpertisePython ProficiencyGPU Kernel OptimizationNVIDIA TensorRT Familiarity
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Deep LearningGenerative AIPerformance AnalysisModel OptimizationPythonPyTorchCUDAAlgorithmsMachine Learning FrameworksHigh-Performance Computing
Soft Skills
CollaborationCommunicationIndependent WorkProblem-Solving
Tools & Technologies
NVIDIA AI Software StackTorch 2.0HuggingFaceTensorRTTRT-LLMTRT Model OptimizerTorchDynamoTorch.exportTorch.compileCUTLASS
Industry Keywords
GPU ArchitectureModel ShardingTensor ParallelismSequence ParallelismKV-CachingEnd-to-End PerformanceModular Software DesignScalable Software PlatformFast-Paced EnvironmentResearch Experience
Tech Stack
Tools & technologiesPythonPyTorch
About the role
Key responsibilities & impact- Train, develop, and deploy state-of-the-art generative AI models such as LLMs and diffusion models using NVIDIA's AI software stack
- Leverage and build upon the Torch 2.0 ecosystem, including TorchDynamo, torch.export, and torch.compile, to analyze and extract standardized model graph representations from arbitrary Torch models
- Develop high-performance inference optimization techniques, including automated model sharding, tensor parallelism, sequence parallelism, and efficient attention kernels with KV-caching
- Collaborate with teams across NVIDIA to integrate performant kernel implementations into the automated deployment solution
- Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities
- Continuously innovate inference performance for NVIDIA's TRT, TRT-LLM, and TRT Model Optimizer software solutions
- Architect and design a modular, scalable software platform supporting broad model coverage and optimization techniques
Requirements
What you’ll need- Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field
- 8+ years of relevant work or research experience in Deep Learning
- Excellent software design skills, including debugging, performance analysis, and test design
- Strong proficiency in Python, PyTorch, and related ML tools (e.g. HuggingFace)
- Strong algorithms and programming fundamentals
- Good written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment
- Contributions to PyTorch, JAX, or other Machine Learning Frameworks
- Knowledge of GPU architecture and compilation stack, and capability of understanding and debugging end-to-end performance
- Familiarity with NVIDIA's deep learning SDKs such as TensorRT
- Prior experience in writing high-performance GPU kernels for machine learning workloads in frameworks such as CUDA, CUTLASS, or Triton
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package
- Equity