FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Compute Engineer
BreakmarkHR. Define lifecycle standards for NVIDIA HGX platforms and validated baselines for firmware, BIOS, BMC, OS, drivers, CUDA, and NCCL .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in NVIDIA GPU infrastructure, including HGX platforms, with a strong focus on burn-in and qualification processes for multi-node GPU clusters. Proficient in root-cause analysis, performance troubleshooting, and defining validation strategies across hardware and software domains.
Highest-signal resume keywords
NVIDIA GPU InfrastructureBurn-In And Qualification DesignLinux Systems AdministrationRoot-Cause AnalysisPerformance Troubleshooting
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
NVIDIA DriversCUDANCCLDCGMGPU Platform ToolingFleet-Level AnalysisFirmware BaselinesBIOS BaselinesOS BaselinesMulti-Node GPU Communication
Soft Skills
Clear CommunicationProblem-SolvingAdaptability
Tools & Technologies
NVIDIA HGXPCIeNVLinkNVSwitchThermal ManagementPower ManagementHigh-Performance AI Fabrics
Industry Keywords
GPU ClustersDistributed TrainingIntra-Node CommunicationValidation StrategiesDeployment Environments
Tech Stack
Tools & technologiesLinuxNode.js
About the role
Key responsibilities & impact- Define lifecycle standards for NVIDIA HGX platforms and validated baselines for firmware, BIOS, BMC, OS, drivers, CUDA, and NCCL
- Own compatibility requirements and model GPU/CPU/memory/PCIe/NVLink/NVSwitch/NIC/storage relationships
- Decide testing required for platform changes before production
- Design qualification and burn-in strategies for new clusters, GPU platforms, repaired nodes, and major changes
- Analyze fleet-wide distributions and outliers to identify isolated versus systemic hardware, firmware, software, topology, power, thermal, or test-configuration issues
- Interpret NVIDIA diagnostics and telemetry across GPU, NVLink, NVSwitch, PCIe, memory, network, thermal, and power signals
- Lead root-cause analysis for intermittent, multi-node, and fleet-wide issues
- Define representative training and inference workloads and reproducible baselines for throughput, scaling, latency, and numerical behavior
- Determine whether performance degradation originates in compute, communication, storage, scheduling, or configuration
- Integrate work across Network, Storage, Platform/DevOps, and Facilities domains
- Route issues to appropriate specialists with supporting evidence
- Define health, readiness, quarantine, and return-to-service gates
- Guide automation teams in building reproducible workflows and interpret validation results
- Translate customer requirements into validation strategies and classify platforms as ready, degraded, conditional, or unsuitable
- Define remediation and re-test scope and communicate risks to engineering, vendors, customers, and leadership
Requirements
What you’ll need- Strong hands-on experience with NVIDIA GPU infrastructure, including HGX or comparable large GPU server platforms
- Demonstrated experience designing, executing, or leading burn-in and qualification across multi-node GPU clusters on H100, H200, B200, B300, or a comparable platform
- Experience analyzing fleet-level qualification results, performance distributions, and hardware/software failures
- Strong Linux systems administration and troubleshooting
- Experience with NVIDIA drivers, CUDA, NCCL, DCGM, and GPU platform tooling
- Experience diagnosing GPU, NVLink, NVSwitch, PCIe, thermal, power, or firmware problems
- Experience validating intra-node and multi-node GPU communication, with a working understanding of RDMA, GPUDirect RDMA, and high-performance AI fabrics
- Experience deploying, provisioning, or operationalizing GPU clusters
- Experience defining firmware, BIOS, OS, and software baselines, as well as health, quarantine, readiness, and return-to-service criteria
- Experience interpreting distributed training or inference workload behavior
- Ability to troubleshoot performance across GPU, host, network, storage, and software boundaries
- Ability to define requirements for automation teams and interpret automated validation outputs
- Strong root-cause analysis and problem-solving skills
- Clear communication of technical findings, risk, and remediation requirements
- Comfortable in fast-moving deployment, commissioning, and production environments
- English and Spanish proficiency levels are requested in the application form
- Must currently live in Medellín or be willing to relocate there