Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Cloud Infrastructure, DevOps Solutions Architect

NVIDIA

. Own full-solution validation on partner software stacks, including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in .

Posted 9/18/2026full-timeRemote • France, United Kingdom, Spain, GermanySeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing scalable cloud environments, particularly with Kubernetes and HPC/AI clusters, while providing consultative guidance and hands-on troubleshooting across diverse technology stacks. Proficient in automation engineering, fault detection, and production stability at fleet scale.

Highest-signal resume keywords
Kubernetes ExperienceHPC/AI Cluster ManagementPython and Bash ScriptingNetworking FundamentalsInfrastructure-as-Code Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesHPC Cluster ManagementPythonBash ScriptingAnsibleTerraformNVIDIA GPU ManagementLinuxStorage SolutionsCI/CD Pipelines
Soft Skills
Consultative GuidanceTechnical LeadershipKnowledge Transfer
Tools & Technologies
GrafanaLokiPrometheusCUDA ToolkitSlurmKubeVirtNVIDIA Base Command ManagerDCGMGitOpsObservability Stacks
Industry Keywords
Cloud EnvironmentsData Centre ArchitecturesMulti-Tenant EstatesAI-Native SchedulingRDMA-Based Fabrics

Tech Stack

Tools & technologies
AnsibleCloudGrafanaKubernetesLinuxMicroservicesNode.jsPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Own full-solution validation on partner software stacks, including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in
  • Minimise time from cluster handover to first production workload by coordinating hardware bring-up, managed-service intake, and partner operations teams
  • Own Day 2 production stability at fleet scale, including monitoring, logging, workload orchestration, fault detection and remediation, preventive maintenance, and firmware and field-notice rollout campaigns
  • Assess customer environments and operate heterogeneous open platforms including Kubernetes, KubeVirt, Slurm, and GPU-aware schedulers
  • Integrate enterprise-grade networking and storage and enable third-party ISV workloads
  • Provide consultative guidance and hands-on troubleshooting across bare metal, operating systems, software stacks, container platforms, networking, and storage
  • Support R&D, proofs of concept, and proofs of value validating new features, architectures, and upgrade approaches
  • Act as technical leader for assigned accounts
  • Run structured knowledge transfer and enablement
  • Produce runbooks, onboarding materials, and best-practice guides for partner teams

Requirements

What you’ll need
  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience
  • 8+ years in managing scalable cloud environments and automation engineering roles
  • Proven understanding of networking fundamentals and data centre architectures
  • Hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure, including deployment, driver and CUDA toolkit management, optimisation, workload profiling, and troubleshooting
  • Extensive Kubernetes experience for container orchestration, resource scheduling, and scaling in GPU-accelerated and HPC environments
  • Experience with scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised multi-tenant estates such as KubeVirt
  • Deep knowledge of Linux, including RedHat and Ubuntu, OS-level security, and protocols
  • Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and Kubernetes storage technologies
  • Proficiency in Python and Bash scripting
  • Experience with configuration management and Infrastructure-as-Code tools such as Ansible and Terraform
  • Experience with GitOps-based cluster lifecycle and upgrade management for large fleets
  • Experience with observability stacks such as Grafana, Loki, and Prometheus
  • Ability to measure and improve MTBI and job goodput on large GPU clusters
  • Experience with fault detection, drain and remediation workflows, SLO/error-budget definition, and post-incident review
  • Strong consultative background leading architectural reviews and presenting to executive stakeholders
  • Knowledge of CI/CD pipelines and container-based microservices architectures
  • Experience with NVIDIA GPU and Network Operators and NVIDIA Base Command Manager
  • Familiarity with DCGM, XID diagnostics, node-level health agents, and fleet-wide reliability intelligence
  • Expertise in AI-native scheduling and inference frameworks on Kubernetes, such as KAI, Grove, Dynamo, and NVIDIA Cloud Functions
  • Background with RDMA-based fabrics such as InfiniBand or RoCE
  • Exposure to Cumulus Linux, SONiC, Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations is a strong plus

Benefits

Comp & perks
  • NVIDIA products and work on next-generation AI/HPC systems
  • Exposure to large-scale infrastructure projects and advanced GPU/HPC technologies
  • Customer, partner, and cross-functional collaboration
  • Technical leadership, knowledge transfer, and enablement opportunities