FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

GPU Cluster Architect
Nebius Group. Architect scalable GPU cluster topologies including compute nodes, InfiniBand/Ethernet interconnects, storage, and control planes .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting scalable GPU cluster topologies and optimizing performance for AI/ML workloads. Proficient in systems architecture, networking, and automation scripting to enhance operational efficiency.
Highest-signal resume keywords
GPU Cluster ArchitectureAI/ML Workload AnalysisInfiniBand HDR/NDRPython ScriptingSystems Architecture
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GPU ArchitectureCluster DesignHPC InterconnectsAutomation ScriptingTelemetry Pipelines
Tools & Technologies
InfiniBandRoCENVIDIAAMD
Industry Keywords
Latency OptimizationBandwidth TradeoffsData-Center EngineeringHardware Reliability
Tech Stack
Tools & technologiesPythonGo
About the role
Key responsibilities & impact- Architect scalable GPU cluster topologies including compute nodes, InfiniBand/Ethernet interconnects, storage, and control planes
- Analyze AI/ML workloads such as LLM training and inference to inform latency, bandwidth, and GPU-density tradeoffs
- Align with network architects and validate low-latency, high-throughput interconnects such as InfiniBand HDR/NDR and RoCEv2 at POD and data-center scale
- Work with storage teams to optimize performance for training datasets and checkpointing
- Analyze monitoring-system signals to detect design flaws
- Partner with site reliability, networking, storage, and data-center engineering teams to operationalize and scale the architecture
Requirements
What you’ll need- 5+ years of experience designing clusters
- Deep understanding of modern GPU architecture, including NVIDIA and AMD
- Experience with HPC interconnects, including InfiniBand and RoCE
- Solid background in systems architecture, networking, and hardware reliability
- Experience scripting for automation and telemetry pipelines using Python, Go, or similar
- Must be authorized to work in the United States and provide proof of employment eligibility
Benefits
Comp & perks- 100% company-paid medical, dental, and vision coverage for employees and families
- 401(k) plan with up to 4% company match and immediate vesting
- 20 weeks paid parental leave for primary caregivers
- 12 weeks paid parental leave for secondary caregivers
- Remote work reimbursement up to $85/month for mobile and internet
- Company-paid short-term, long-term, and life insurance coverage
- Equity in the form of RSUs may be available at certain salary grades
- Opportunities for professional growth within Nebius
- Hybrid working arrangements
- Competitive salary and comprehensive benefits package
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams