FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in leading engineering teams and managing infrastructure platforms, with a strong focus on Slurm and Kubernetes operations. Proficient in translating technical requirements into actionable specifications while ensuring high standards in performance and reliability.
Highest-signal resume keywords
Slurm ManagementKubernetes OperationsHPC Infrastructure ExperiencePeople ManagementLinux Systems Knowledge
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SlurmctldSlurmdbdNVIDIA DriverCUDA LifecycleInfiniBandRDMAJavaScriptNode.jsPrometheusGrafana
Soft Skills
Written CommunicationTeam LeadershipTechnical Judgment
Tools & Technologies
KubernetesFabric ManagerNVSwitchDCGMLustreNFSVM/Container IsolationGitOpsCiliumCluster API
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree in Engineering
Industry Keywords
HPCGPU Training ClusterResearch Computing ServicesMulti-Tenant IaaS/PaaSDistributed Systems
Tech Stack
Tools & technologiesBootstrapCloudGrafanaJavaScriptKubernetesLinuxNFSNode.jsPrometheus
About the role
Key responsibilities & impact- Own Cosmic AC platform architecture end to end through architecture proposals, high-level and low-level designs, reviews, and baseline maintenance
- Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation
- Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
- Design, build, and operate a managed Slurm service for research users
- Own Slurm controller/accounting, partitions, login nodes, node onboarding, driver/CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation
- Own Kubernetes cluster bootstrap and lifecycle on partner bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations
- Define managed inference architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
- Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and sustainable on-call operations
- Serve as primary technical interface to infrastructure partners and vendors
- Translate requirements into written specifications and acceptance tests, run escalations to closure, and support capacity planning and hardware sourcing
- Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
- Complete the platform team and set the technical bar for new engineers
Requirements
What you’ll need- Eight or more years of hands-on engineering experience
- At least three years leading teams that build and operate infrastructure platforms other teams depend on
- Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
- Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
- Experience operating an HPC or GPU training cluster for a research population is ideally required
- Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
- Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues
- Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
- Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
- Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
- Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
- Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
- Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing services
- People management across time zones, cross-track review, written architecture decisions, and technical judgment with partners and executives
- Excellent written and spoken English
- Fully remote and based between UTC and UTC+5:30
- Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
- Desirable: modern serving stacks including vLLM, SGLang, and TensorRT-LLM
- Desirable: VM/container isolation and confidential computing technologies
- Desirable: Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, and GitOps
- Desirable: GPU cloud, HPC centre, AI lab platform, peer-to-peer, or distributed-systems experience
- Desirable: hardware-provider relationship and written contract/acceptance-test experience
Benefits
Comp & perks- Fully remote work arrangement
- Occasional travel to partner sites and team events
