Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Software Engineer, DGX Cloud AI Infrastructure

NVIDIA

. Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads .

Posted 9/30/2026full-timeRemote • United StatesMid-LevelSenior💰 $108,000 - $178,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and debugging large-scale AI clusters and distributed systems, with strong proficiency in Python and C/C++. Capable of delivering data-driven recommendations and building resilience and failure-attribution systems for datacenter-scale infrastructure.

Highest-signal resume keywords
AI Cluster DevelopmentPython ProgrammingC/C++ ProgrammingCUDA-Aware Distributed ExecutionBenchmarking and Performance Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI Workload DebuggingMulti-GPU WorkloadsNCCLRDMA Software StackPerformance Jitter DiagnosisBenchmark Harness DevelopmentFailure-Attribution SystemsRoot-Cause AnalysisAutomation WorkflowsCluster Qualification Tooling
Soft Skills
Analytical SkillsCommunication SkillsCollaborative Approach
Tools & Technologies
PyTorchNeMoMegatronTensorRT-LLMInfiniBandRoCE
Industry Keywords
AIHPCDistributed SystemsDatacenter InfrastructureContainerized Environments

Tech Stack

Tools & technologies
Distributed SystemsNode.jsPythonPyTorchC++

About the role

Key responsibilities & impact
  • Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads
  • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks
  • Perform root-cause analysis of failures in large distributed environments
  • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster
  • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms
  • Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams
  • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization

Requirements

What you’ll need
  • Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience)
  • Experience developing software for AI, HPC, or systems-level applications
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
  • Background with debugging and scaling distributed systems
  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
  • Experience operating workloads in scheduled, containerized cluster environments
  • Excellent analytical, debugging, and communication skills, and a collaborative approach across teams
  • Strong Python and C/C++ programming skills
  • Hands-on experience with NCCL and CUDA-aware distributed execution
  • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging
  • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
  • Experience diagnosing performance jitter
  • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score