Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior System Software Engineer – Dynamo-Triton Inference Server

NVIDIA

. Develop GPU-accelerated AI inference serving software .

Posted 9/18/2026full-timeSanta Clara • California • United StatesSenior💰 $224,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing GPU-accelerated AI inference serving software, with a strong focus on deep learning frameworks and high-scale distributed systems. Proficient in Rust and C++, with a solid understanding of performance optimization and software design principles.

Highest-signal resume keywords
GPU-Accelerated AI Inference DevelopmentRust ProgrammingC++ ProgrammingDeep Learning FrameworksDistributed Systems Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Deep Learning Software DevelopmentPerformance AnalysisDebuggingTest DesignHigh-Scale Distributed SystemsML SystemsGPU Memory ManagementCache ManagementAI FrameworksSoftware Design
Soft Skills
Strong Communication SkillsAgile Team Collaboration
Tools & Technologies
Triton Inference ServerNVIDIA DynamoTensorRTPyTorchONNXOpenVINOVLLMTRT-LLMGitHubBug Tracking
Certifications & Qualifications
MS or PhD in Computer Science
Industry Keywords
AI InferenceOpen Source SoftwareHigh-Performance NetworkingLarge Language ModelsProduction Server Environments

Tech Stack

Tools & technologies
CloudDistributed SystemsOpen SourcePythonPyTorchRustC++

About the role

Key responsibilities & impact
  • Develop GPU-accelerated AI inference serving software
  • Contribute to feature development and drive broad customer adoption
  • Drive convergence of the Triton Inference Server and NVIDIA Dynamo stacks to establish a unified, high-performance inference platform
  • Ensure feature parity while serving Large Language Model (LLM) and non-LLM workloads
  • Participate in the open source deep learning software engineering community
  • Build robust software for production server or cloud environments
  • Optimize and balance prediction throughput and latency
  • Develop and adopt next-generation inference technologies

Requirements

What you’ll need
  • MS or PhD in Computer Science or relevant field (or equivalent experience)
  • 12+ years of professional experience working on deep learning software
  • Excellent Rust and C++ skills
  • Familiarity with Python
  • Strong programming and software design skills, including debugging, performance analysis, and test design
  • Experience with high-scale distributed systems and ML systems
  • Strong communication skills and ability to work in a fast-paced, agile team environment
  • Prior experience with AI frameworks and engines such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM
  • Knowledge of GPU memory management, cache management, or high-performance networking
  • Experience with distributed systems programming
  • Experience contributing to a large open source project, including GitHub, bug tracking, branching and merging code, OSS licensing issues, and handling patches

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score