Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Software Engineer – Cluster Networking

NVIDIA

. Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale .

Posted 9/15/2026full-timeRemote • California • United StatesSenior💰 $184,000 - $287,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Expertise in Kubernetes Networking Architecture and CNI Standards, with a strong focus on designing and operating large-scale networking topologies. Proficient in debugging distributed network issues and collaborating effectively with cross-functional teams.

Highest-signal resume keywords
Kubernetes Networking ArchitectureCNI StandardsGo ProgrammingLinux Networking FundamentalsHigh-Performance Fabrics

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesCNICalicoTailscaleWireGuardGoPythonCPacket CaptureDebugging Distributed Networks
Soft Skills
Clear CommunicationCollaboration Across Time Zones
Tools & Technologies
EnvoyCloudflareSlurmInfiniBandRoCERDMA
Industry Keywords
AIHPCNetworking TopologiesOverlay NetworkLoad Balancers

Tech Stack

Tools & technologies
CloudKubernetesLinuxNode.jsPythonGo

About the role

Key responsibilities & impact
  • Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale
  • Design, operate and scale the overlay network - CNI, mesh and VPN topologies (Tailscale, WireGuard), and the gateways that connect control and data planes
  • Design, operate and scale the L7 gateways/load balancers/tunnels (Envoy, Cloudflare)
  • Find and eliminate scale ceilings: packet loss under load, control-plane saturation, IP address management exhaustion, and failure modes above a few thousand nodes
  • Build scale-test environments and validation suites to catch networking regressions before production
  • Diagnose hard, ambiguous problems across the stack, tracing symptoms in Slurm or training jobs to networking root causes
  • Partner with cloud and neocloud providers on network topology, requirements and capabilities for new clusters
  • Provide senior technical judgement to a distributed team in the Cluster Networking domain

Requirements

What you’ll need
  • BS/MS in Computer Science, Electrical Engineering or a related field, or equivalent experience
  • 6+ years of professional experience in systems, network or infrastructure software engineering
  • Deep command of Kubernetes networking architecture and CNI standards, with production experience operating Calico strongly preferred
  • Proficiency designing and maintaining modern mesh and VPN networking topologies - Tailscale, WireGuard or equivalent
  • Strong Linux networking fundamentals: routing, netfilter and iptables/nftables, packet marking, network namespaces, and how these interact with container runtimes
  • Demonstrated ability to debug distributed network problems at scale - packet capture, tracing, and correlating behaviour across many hosts to find a single root cause
  • Proficiency in Go, Python, C or a comparable systems language
  • Clear written and verbal communication, and the ability to work effectively with engineers across multiple time zones
  • Direct experience architecting and operating massive-scale Kubernetes topologies across thousands of concurrent nodes
  • Experience with high-performance fabrics in AI or HPC environments - InfiniBand, RoCE, or RDMA over converged networks
  • Upstream contributions to Calico, Cilium, Tailscale, or Kubernetes networking SIGs
  • Experience operating networking across multiple public clouds and on-premises environments simultaneously
  • Familiarity with Slurm or other HPC schedulers running on Kubernetes

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score