FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesDistributed SystemsKubernetesLinux
About the role
Key responsibilities & impact- Build and improve Kubernetes infrastructure for provisioning, networking, storage, and service deployment across providers and regions
- Design isolation, failover, and recovery mechanisms to reduce the impact of hardware and infrastructure failures
- Automate capacity expansion, deployments, and maintenance
- Develop observability and diagnostics to reveal bottlenecks and surface failures
- Respond to incidents, identify root causes, and improve systems based on findings
- Improve platform performance, utilization, security, and reliability as inference demand grows
- Work with infrastructure, platform, and inference engineers in a flat organization
Requirements
What you’ll need- Experience building and operating production infrastructure or distributed systems, with real ownership of reliability
- Strong Linux fundamentals
- Practical knowledge of networking, storage, and containers
- Hands-on experience running Kubernetes in production
- Ability to write maintainable software and automation for infrastructure problems
- Systematic debugging across application, cluster, network, and hardware boundaries
- Good judgment about prioritization, simplification, and reliability
- Initiative to take problems from investigation through implementation
- Experience with multi-region, multi-provider, or bare-metal infrastructure (nice to have)
- Familiarity with GPUs, model serving, vLLM, or SGLang (nice to have)
- Experience with infrastructure as code, CI/CD, observability, or automated recovery (nice to have)
- Experience building highly available services, multi-tenant platforms, or distributed data systems (nice to have)
Benefits
Comp & perks- Ownership and reach to shape the company’s scaling infrastructure
- Opportunity to work directly with infrastructure, platform, and inference engineers
- Meaningful architecture ownership
- Ship improvements directly into production
- Work close to hardware and deep in distributed systems
- Help build the foundation for the next stage of AI infrastructure
