FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and debugging large-scale AI clusters and distributed systems, with strong proficiency in Python and C/C++. Capable of delivering data-driven recommendations and building resilience and failure-attribution systems for datacenter-scale infrastructure.
Highest-signal resume keywords
AI Cluster DevelopmentPython ProgrammingC/C++ ProgrammingCUDA-Aware Distributed ExecutionBenchmarking and Performance Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI Workload DebuggingMulti-GPU WorkloadsNCCLRDMA Software StackPerformance Jitter DiagnosisBenchmark Harness DevelopmentFailure-Attribution SystemsRoot-Cause AnalysisAutomation WorkflowsCluster Qualification Tooling
Soft Skills
Analytical SkillsCommunication SkillsCollaborative Approach
Tools & Technologies
PyTorchNeMoMegatronTensorRT-LLMInfiniBandRoCE
Industry Keywords
AIHPCDistributed SystemsDatacenter InfrastructureContainerized Environments
Tech Stack
Tools & technologiesDistributed SystemsNode.jsPythonPyTorchC++
About the role
Key responsibilities & impact- Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads
- Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks
- Perform root-cause analysis of failures in large distributed environments
- Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster
- Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms
- Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams
- Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization
Requirements
What you’ll need- Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience)
- Experience developing software for AI, HPC, or systems-level applications
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
- Background with debugging and scaling distributed systems
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
- Experience operating workloads in scheduled, containerized cluster environments
- Excellent analytical, debugging, and communication skills, and a collaborative approach across teams
- Strong Python and C/C++ programming skills
- Hands-on experience with NCCL and CUDA-aware distributed execution
- Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging
- Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
- Experience diagnosing performance jitter
- Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score
