Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
BreakmarkHR

AI Compute Engineer

BreakmarkHR

. Define lifecycle standards for NVIDIA HGX platforms and validated baselines for firmware, BIOS, BMC, OS, drivers, CUDA, and NCCL .

Posted 10/5/2026full-timeMedellín • ColombiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in NVIDIA GPU infrastructure, including HGX platforms, with a strong focus on burn-in and qualification processes for multi-node GPU clusters. Proficient in root-cause analysis, performance troubleshooting, and defining validation strategies across hardware and software domains.

Highest-signal resume keywords
NVIDIA GPU InfrastructureBurn-In And Qualification DesignLinux Systems AdministrationRoot-Cause AnalysisPerformance Troubleshooting

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
NVIDIA DriversCUDANCCLDCGMGPU Platform ToolingFleet-Level AnalysisFirmware BaselinesBIOS BaselinesOS BaselinesMulti-Node GPU Communication
Soft Skills
Clear CommunicationProblem-SolvingAdaptability
Tools & Technologies
NVIDIA HGXPCIeNVLinkNVSwitchThermal ManagementPower ManagementHigh-Performance AI Fabrics
Industry Keywords
GPU ClustersDistributed TrainingIntra-Node CommunicationValidation StrategiesDeployment Environments

Tech Stack

Tools & technologies
LinuxNode.js

About the role

Key responsibilities & impact
  • Define lifecycle standards for NVIDIA HGX platforms and validated baselines for firmware, BIOS, BMC, OS, drivers, CUDA, and NCCL
  • Own compatibility requirements and model GPU/CPU/memory/PCIe/NVLink/NVSwitch/NIC/storage relationships
  • Decide testing required for platform changes before production
  • Design qualification and burn-in strategies for new clusters, GPU platforms, repaired nodes, and major changes
  • Analyze fleet-wide distributions and outliers to identify isolated versus systemic hardware, firmware, software, topology, power, thermal, or test-configuration issues
  • Interpret NVIDIA diagnostics and telemetry across GPU, NVLink, NVSwitch, PCIe, memory, network, thermal, and power signals
  • Lead root-cause analysis for intermittent, multi-node, and fleet-wide issues
  • Define representative training and inference workloads and reproducible baselines for throughput, scaling, latency, and numerical behavior
  • Determine whether performance degradation originates in compute, communication, storage, scheduling, or configuration
  • Integrate work across Network, Storage, Platform/DevOps, and Facilities domains
  • Route issues to appropriate specialists with supporting evidence
  • Define health, readiness, quarantine, and return-to-service gates
  • Guide automation teams in building reproducible workflows and interpret validation results
  • Translate customer requirements into validation strategies and classify platforms as ready, degraded, conditional, or unsuitable
  • Define remediation and re-test scope and communicate risks to engineering, vendors, customers, and leadership

Requirements

What you’ll need
  • Strong hands-on experience with NVIDIA GPU infrastructure, including HGX or comparable large GPU server platforms
  • Demonstrated experience designing, executing, or leading burn-in and qualification across multi-node GPU clusters on H100, H200, B200, B300, or a comparable platform
  • Experience analyzing fleet-level qualification results, performance distributions, and hardware/software failures
  • Strong Linux systems administration and troubleshooting
  • Experience with NVIDIA drivers, CUDA, NCCL, DCGM, and GPU platform tooling
  • Experience diagnosing GPU, NVLink, NVSwitch, PCIe, thermal, power, or firmware problems
  • Experience validating intra-node and multi-node GPU communication, with a working understanding of RDMA, GPUDirect RDMA, and high-performance AI fabrics
  • Experience deploying, provisioning, or operationalizing GPU clusters
  • Experience defining firmware, BIOS, OS, and software baselines, as well as health, quarantine, readiness, and return-to-service criteria
  • Experience interpreting distributed training or inference workload behavior
  • Ability to troubleshoot performance across GPU, host, network, storage, and software boundaries
  • Ability to define requirements for automation teams and interpret automated validation outputs
  • Strong root-cause analysis and problem-solving skills
  • Clear communication of technical findings, risk, and remediation requirements
  • Comfortable in fast-moving deployment, commissioning, and production environments
  • English and Spanish proficiency levels are requested in the application form
  • Must currently live in Medellín or be willing to relocate there