FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in platform architecture, leading distributed engineering teams, and managing complex infrastructure operations. Proficient in Slurm, Kubernetes, and GPU cluster management, with a strong focus on performance metrics and incident response.
Highest-signal resume keywords
Slurm ManagementKubernetes OperationsGPU Cluster ManagementPeople ManagementLinux Systems Knowledge
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SlurmKubernetesJavaScriptNode.jsHPC StorageInfiniBandPerformance TuningMulti-Tenancy DesignIncident ResponseMetrics and Logging
Soft Skills
LeadershipCommunicationTeam ManagementTechnical JudgmentWritten Specifications
Tools & Technologies
PrometheusGrafanaLustreNFSCUDAFabric ManagerDCGMNVSwitchKata ContainersQEMU/KVM
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree in Engineering
Industry Keywords
Infrastructure PlatformsResearch ComputingMulti-Tenant IaaS/PaaSDistributed SystemsCloud Computing
Tech Stack
Tools & technologiesBootstrapCloudDistributed SystemsGrafanaJavaScriptKubernetesLinuxNFSNode.jsPrometheus
About the role
Key responsibilities & impact- Own the end-to-end platform architecture through proposals, high-level and low-level designs, reviews, and baseline maintenance
- Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation
- Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
- Design, build, and operate a managed Slurm service for research users
- Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health detection, autohealing, storage visibility, identity, and isolation
- Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU and Network Operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement
- Design managed inference architecture, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
- Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model
- Serve as primary technical interface to infrastructure partners and vendors
- Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
- Work with research, model-training, and product teams to translate workloads into platform requirements and broker scarce capacity
- Complete the platform team and set the technical bar for new engineers
Requirements
What you’ll need- Eight or more years of hands-on engineering experience
- At least three years leading teams that build and operate infrastructure platforms other teams depend on
- Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
- Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running
- Experience operating an HPC or GPU training cluster for a research population is ideally preferred
- Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
- Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
- Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
- Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
- Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
- Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
- Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control plane, CLI, and worker services
- Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing service
- People management across time zones, cross-track review, written architecture decisions, and technical judgment with partners and executives
- Excellent written and spoken English
- Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers
- Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM
- Desirable experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships
Benefits
Comp & perks- Fully remote work
- Occasional travel to partner sites and team events
