FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetes
About the role
Key responsibilities & impact- Lead the Capacity Engineering team and establish its 12–18 month functional and technical strategy, roadmap and success measures
- Develop CPU and GPU forecasts and supply plans accounting for workload demand, delivery constraints, cost and reliability requirements
- Guide capacity requests, reservations, entitlements, allocation policy and infrastructure-data systems
- Improve utilization and efficiency across Kubernetes-based shared compute platforms, GPU infrastructure and other cloud environments
- Establish metrics and operating practices for capacity risk, allocation, infrastructure cost, realized savings and system health
- Partner with Infrastructure, SRE, Finance, product and platform teams and cloud providers on capacity and infrastructure-investment decisions
- Plan team structure and headcount, recruit senior talent, develop technical leaders and sustain an inclusive, accountable, high-performing team
- Use AI to accelerate capacity analysis, forecasting, operational investigations and repeatable workflows while applying judgment and verification
Requirements
What you’ll need- Bachelor’s degree in computer science, a related field or equivalent experience
- Experience managing and developing teams building or operating large-scale cloud infrastructure, distributed systems, compute platforms or resource-management systems
- Experience defining and delivering a 12–18 month functional and technical strategy through senior technical leaders across a broad infrastructure portfolio
- Technical experience in public-cloud infrastructure, highly available distributed systems, Kubernetes or shared-compute platforms and capacity-management systems
- Experience leading cross-functional infrastructure programs and communicating availability, cost and capacity tradeoffs to Engineering, Finance, SRE and cloud-provider stakeholders
- Demonstrated ability to use AI to improve speed and quality in day-to-day workflow
- Critical evaluation and verification of AI-assisted work, including testing, source-checking, data validation and peer review
- High integrity and ownership, including protecting sensitive data, avoiding over-reliance on AI and remaining accountable for final decisions and deliverables
Benefits
Comp & perks- Equity eligibility
- Flexible working model
- In-person collaboration 1–2 times per quarter
- Equal opportunity employment
- Medical or religious accommodation support during the application process
