FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing network solutions for GPU clouds and AI factories, with a strong focus on InfiniBand and RoCEv2 Ethernet technologies. Proficient in Kubernetes networking, data center architecture, and customer engagement in technical workshops and presentations.
Highest-signal resume keywords
InfiniBand ExpertiseRoCEv2 Ethernet KnowledgeKubernetes NetworkingData Center Network DesignProgramming in Python or Go
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Network Architecture DesignGPU Cluster InterconnectsMulti-Tenant IsolationLinux Host NetworkingKubeVirt NetworkingData Center NetworkingNCCL TestingAutomation with AnsibleInfrastructure as CodeNetwork Policy Implementation
Soft Skills
Excellent CommunicationCustomer-Facing ExperiencePresentation SkillsCollaborationTechnical Workshop Facilitation
Tools & Technologies
KubernetesGitTerraformBashHelmNVIDIA Network OperatorRDMA Device PluginsIproute2EthtoolDevlink
Industry Keywords
HPC NetworkingCloud InfrastructureAI WorkloadsNetwork AutomationSDNMulti-Cloud ConnectivityVPNDNS PlanningGPU Cloud DeploymentsOpen-Source Contributions
Tech Stack
Tools & technologiesAnsibleCloudDNSKubernetesLinuxPythonTerraformGo
About the role
Key responsibilities & impact- Own network solutioning for k0rdent AI, Mirantis's platform for building and operating GPU clouds and AI factories
- Design large GPU cluster interconnects, tenant isolation and workload connectivity from containers, virtual machines and bare-metal nodes
- Design and publish network reference architectures and solution designs from single-rack to multi-thousand-GPU clusters
- Define InfiniBand and RoCEv2 Ethernet compute fabrics, front-end, storage and management networks
- Specify multi-tenant isolation using InfiniBand partitions, VRFs, EVPN-VXLAN and Kubernetes network policy
- Document scale limits and trade-offs involving cost, performance, operability and vendor lock-in
- Define Linux host networking for NICs, DPUs, SuperNICs, PCI passthrough, SR-IOV, IOMMU, NUMA and GPU-NIC affinity
- Design Kubernetes networking for AI workloads using CNIs, Multus, SR-IOV, RDMA device plugins, NVIDIA Network Operator and DRA
- Design KubeVirt networking for GPU and RDMA traffic
- Research emerging AI networking technologies and standards, prototype alternative designs and benchmark them
- Publish research notes, design proposals, papers, blog posts and talks
- Build and run proofs of concept with customers, partners and on Mirantis hardware or customer sites
- Write automation and tooling using Python, Go, Bash, Ansible, Helm, Kubernetes manifests and Terraform
- Validate and benchmark fabrics and host configurations using NCCL tests, perftest and ib_write_bw
- Act as network subject matter expert in customer discovery, design reviews and architecture workshops
- Collaborate with hardware and networking partners and feed research into Product, Engineering and the k0rdent AI roadmap
- Present at industry events, webinars and partner summits and run technical workshops
- Write reference architectures, solution briefs, blog posts and internal enablement content
Requirements
What you’ll need- Bachelor's degree in Computer Science, Electrical Engineering, Telecommunications or a related field, or equivalent practical experience
- 8+ years in network engineering or network architecture
- At least 3 years in data center, HPC or cloud infrastructure networking
- Customer-facing experience as a solutions architect, pre-sales engineer, consultant or technical lead
- Experience with data center network design, including spine-leaf and Clos topologies, rail-optimised GPU fabrics, oversubscription and ECMP
- Expertise in InfiniBand, including subnet management, partitioning, adaptive routing and NCCL/RDMA traffic
- Expertise in RoCEv2 Ethernet, including PFC, ECN, DCQCN, QoS, MTU and buffer tuning
- Knowledge of BGP, EVPN-VXLAN, VRFs, hybrid and multi-cloud connectivity, VPN, IP address and DNS planning
- Linux networking knowledge, including NICs, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool and devlink
- Kubernetes networking knowledge, including CNI plugins, Multus, SR-IOV and RDMA device plugins, network policy and service exposure
- Experience with KubeVirt VM networking, passthrough and SR-IOV
- Programming or scripting in at least one language; Python or Go preferred
- Comfortable using Git, CI and infrastructure-as-code
- Excellent written and spoken English
- Comfortable presenting to large audiences and running hands-on workshops
- Experience working in an international, distributed company across time zones and work cultures
- Willingness to work some meetings outside standard local hours
- Nice-to-have experience includes NVIDIA networking, GPU cloud/HPC deployments, bare-metal provisioning, network automation, SDN, switch operating systems, GSLB, Gateway API, NVMe-oF, GPUDirect Storage, Mirantis products, open-source contributions, industry standards, research, patents, white papers and additional languages
Benefits
Comp & perks- Remote work arrangement
- Up to 25% travel for customer engagements, partner meetings, lab work and industry events
- Distributed, international team collaboration
- Opportunity to work directly with leading GPU cloud operators, NeoClouds, sovereign clouds and AI-first enterprises
- Opportunity to shape the product narrative and influence go-to-market success
