Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Tether.to

Technical Lead – GPU Infrastructure

Tether.to

. Own Cosmic AC platform architecture end to end through architecture proposals, high-level and low-level designs, reviews, and baseline maintenance .

Posted 9/17/2026full-timeRemote • United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in leading engineering teams and managing infrastructure platforms, with a strong focus on Slurm and Kubernetes operations. Proficient in translating technical requirements into actionable specifications while ensuring high standards in performance and reliability.

Highest-signal resume keywords
Slurm ManagementKubernetes OperationsHPC Infrastructure ExperiencePeople ManagementLinux Systems Knowledge

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SlurmctldSlurmdbdNVIDIA DriverCUDA LifecycleInfiniBandRDMAJavaScriptNode.jsPrometheusGrafana
Soft Skills
Written CommunicationTeam LeadershipTechnical Judgment
Tools & Technologies
KubernetesFabric ManagerNVSwitchDCGMLustreNFSVM/Container IsolationGitOpsCiliumCluster API
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree in Engineering
Industry Keywords
HPCGPU Training ClusterResearch Computing ServicesMulti-Tenant IaaS/PaaSDistributed Systems

Tech Stack

Tools & technologies
BootstrapCloudGrafanaJavaScriptKubernetesLinuxNFSNode.jsPrometheus

About the role

Key responsibilities & impact
  • Own Cosmic AC platform architecture end to end through architecture proposals, high-level and low-level designs, reviews, and baseline maintenance
  • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation
  • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
  • Design, build, and operate a managed Slurm service for research users
  • Own Slurm controller/accounting, partitions, login nodes, node onboarding, driver/CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation
  • Own Kubernetes cluster bootstrap and lifecycle on partner bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations
  • Define managed inference architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
  • Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and sustainable on-call operations
  • Serve as primary technical interface to infrastructure partners and vendors
  • Translate requirements into written specifications and acceptance tests, run escalations to closure, and support capacity planning and hardware sourcing
  • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
  • Complete the platform team and set the technical bar for new engineers

Requirements

What you’ll need
  • Eight or more years of hands-on engineering experience
  • At least three years leading teams that build and operate infrastructure platforms other teams depend on
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
  • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
  • Experience operating an HPC or GPU training cluster for a research population is ideally required
  • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
  • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues
  • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
  • Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
  • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
  • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
  • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
  • Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing services
  • People management across time zones, cross-track review, written architecture decisions, and technical judgment with partners and executives
  • Excellent written and spoken English
  • Fully remote and based between UTC and UTC+5:30
  • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
  • Desirable: modern serving stacks including vLLM, SGLang, and TensorRT-LLM
  • Desirable: VM/container isolation and confidential computing technologies
  • Desirable: Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, and GitOps
  • Desirable: GPU cloud, HPC centre, AI lab platform, peer-to-peer, or distributed-systems experience
  • Desirable: hardware-provider relationship and written contract/acceptance-test experience

Benefits

Comp & perks
  • Fully remote work arrangement
  • Occasional travel to partner sites and team events