Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Nebius Group

GPU Cluster Architect

Nebius Group

. Architect scalable GPU cluster topologies including compute nodes, InfiniBand/Ethernet interconnects, storage, and control planes .

Posted 9/28/2026full-timeRemote • United StatesMid-LevelSenior💰 $184,000 - $318,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting scalable GPU cluster topologies and optimizing performance for AI/ML workloads. Proficient in systems architecture, networking, and automation scripting to enhance operational efficiency.

Highest-signal resume keywords
GPU Cluster ArchitectureAI/ML Workload AnalysisInfiniBand HDR/NDRPython ScriptingSystems Architecture

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GPU ArchitectureCluster DesignHPC InterconnectsAutomation ScriptingTelemetry Pipelines
Tools & Technologies
InfiniBandRoCENVIDIAAMD
Industry Keywords
Latency OptimizationBandwidth TradeoffsData-Center EngineeringHardware Reliability

Tech Stack

Tools & technologies
PythonGo

About the role

Key responsibilities & impact
  • Architect scalable GPU cluster topologies including compute nodes, InfiniBand/Ethernet interconnects, storage, and control planes
  • Analyze AI/ML workloads such as LLM training and inference to inform latency, bandwidth, and GPU-density tradeoffs
  • Align with network architects and validate low-latency, high-throughput interconnects such as InfiniBand HDR/NDR and RoCEv2 at POD and data-center scale
  • Work with storage teams to optimize performance for training datasets and checkpointing
  • Analyze monitoring-system signals to detect design flaws
  • Partner with site reliability, networking, storage, and data-center engineering teams to operationalize and scale the architecture

Requirements

What you’ll need
  • 5+ years of experience designing clusters
  • Deep understanding of modern GPU architecture, including NVIDIA and AMD
  • Experience with HPC interconnects, including InfiniBand and RoCE
  • Solid background in systems architecture, networking, and hardware reliability
  • Experience scripting for automation and telemetry pipelines using Python, Go, or similar
  • Must be authorized to work in the United States and provide proof of employment eligibility

Benefits

Comp & perks
  • 100% company-paid medical, dental, and vision coverage for employees and families
  • 401(k) plan with up to 4% company match and immediate vesting
  • 20 weeks paid parental leave for primary caregivers
  • 12 weeks paid parental leave for secondary caregivers
  • Remote work reimbursement up to $85/month for mobile and internet
  • Company-paid short-term, long-term, and life insurance coverage
  • Equity in the form of RSUs may be available at certain salary grades
  • Opportunities for professional growth within Nebius
  • Hybrid working arrangements
  • Competitive salary and comprehensive benefits package
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams