FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Cloud Infrastructure, DevOps Solutions Architect
NVIDIA. Own full-solution validation on partner software stacks, including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing scalable cloud environments, particularly with Kubernetes and HPC/AI clusters, while providing consultative guidance and hands-on troubleshooting across diverse technology stacks. Proficient in automation engineering, fault detection, and production stability at fleet scale.
Highest-signal resume keywords
Kubernetes ExperienceHPC/AI Cluster ManagementPython and Bash ScriptingNetworking FundamentalsInfrastructure-as-Code Tools
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesHPC Cluster ManagementPythonBash ScriptingAnsibleTerraformNVIDIA GPU ManagementLinuxStorage SolutionsCI/CD Pipelines
Soft Skills
Consultative GuidanceTechnical LeadershipKnowledge Transfer
Tools & Technologies
GrafanaLokiPrometheusCUDA ToolkitSlurmKubeVirtNVIDIA Base Command ManagerDCGMGitOpsObservability Stacks
Industry Keywords
Cloud EnvironmentsData Centre ArchitecturesMulti-Tenant EstatesAI-Native SchedulingRDMA-Based Fabrics
Tech Stack
Tools & technologiesAnsibleCloudGrafanaKubernetesLinuxMicroservicesNode.jsPrometheusPythonTerraform
About the role
Key responsibilities & impact- Own full-solution validation on partner software stacks, including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in
- Minimise time from cluster handover to first production workload by coordinating hardware bring-up, managed-service intake, and partner operations teams
- Own Day 2 production stability at fleet scale, including monitoring, logging, workload orchestration, fault detection and remediation, preventive maintenance, and firmware and field-notice rollout campaigns
- Assess customer environments and operate heterogeneous open platforms including Kubernetes, KubeVirt, Slurm, and GPU-aware schedulers
- Integrate enterprise-grade networking and storage and enable third-party ISV workloads
- Provide consultative guidance and hands-on troubleshooting across bare metal, operating systems, software stacks, container platforms, networking, and storage
- Support R&D, proofs of concept, and proofs of value validating new features, architectures, and upgrade approaches
- Act as technical leader for assigned accounts
- Run structured knowledge transfer and enablement
- Produce runbooks, onboarding materials, and best-practice guides for partner teams
Requirements
What you’ll need- BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience
- 8+ years in managing scalable cloud environments and automation engineering roles
- Proven understanding of networking fundamentals and data centre architectures
- Hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure, including deployment, driver and CUDA toolkit management, optimisation, workload profiling, and troubleshooting
- Extensive Kubernetes experience for container orchestration, resource scheduling, and scaling in GPU-accelerated and HPC environments
- Experience with scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised multi-tenant estates such as KubeVirt
- Deep knowledge of Linux, including RedHat and Ubuntu, OS-level security, and protocols
- Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and Kubernetes storage technologies
- Proficiency in Python and Bash scripting
- Experience with configuration management and Infrastructure-as-Code tools such as Ansible and Terraform
- Experience with GitOps-based cluster lifecycle and upgrade management for large fleets
- Experience with observability stacks such as Grafana, Loki, and Prometheus
- Ability to measure and improve MTBI and job goodput on large GPU clusters
- Experience with fault detection, drain and remediation workflows, SLO/error-budget definition, and post-incident review
- Strong consultative background leading architectural reviews and presenting to executive stakeholders
- Knowledge of CI/CD pipelines and container-based microservices architectures
- Experience with NVIDIA GPU and Network Operators and NVIDIA Base Command Manager
- Familiarity with DCGM, XID diagnostics, node-level health agents, and fleet-wide reliability intelligence
- Expertise in AI-native scheduling and inference frameworks on Kubernetes, such as KAI, Grove, Dynamo, and NVIDIA Cloud Functions
- Background with RDMA-based fabrics such as InfiniBand or RoCE
- Exposure to Cumulus Linux, SONiC, Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations is a strong plus
Benefits
Comp & perks- NVIDIA products and work on next-generation AI/HPC systems
- Exposure to large-scale infrastructure projects and advanced GPU/HPC technologies
- Customer, partner, and cross-functional collaboration
- Technical leadership, knowledge transfer, and enablement opportunities