Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cerebras

AI Fleet Platform Software Engineer

Cerebras

. Build and operate software for managing large fleets of AI clusters .

Posted 10/2/2026full-timeSunnyvale • California • United StatesSeniorLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production software for distributed systems, with strong skills in Go or Python, and a focus on reliability, security, and observability. Capable of leading projects that integrate data and workflows across multiple infrastructure systems while improving tools for incident response and fleet management.

Highest-signal resume keywords
Go ProgrammingPython ProgrammingDistributed Systems DesignKubernetes ManagementEvent Streaming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Control PlanesFleet Management SystemsAPIs DesignAsynchronous Work HandlingWorkflow AutomationTime-Series TelemetryIncident Response Software DevelopmentCapacity Management Software DevelopmentDashboard DevelopmentLinux Systems
Soft Skills
Strong JudgmentLeadershipCollaboration
Tools & Technologies
ContainersKubernetesAI Clusters
Industry Keywords
Infrastructure SystemsOperational PlatformsObservabilitySecurityCluster Health

Tech Stack

Tools & technologies
Distributed SystemsKubernetesLinuxPythonGo

About the role

Key responsibilities & impact
  • Build and operate software for managing large fleets of AI clusters
  • Give operators clear, actionable views of cluster health, capacity, performance, and ongoing issues
  • Develop services and integrations that bring together data and workflows from multiple infrastructure systems
  • Automate repetitive work and improve tools used to investigate incidents and restore service
  • Design systems that remain reliable as the fleet grows and through component and site failures
  • Work with Cluster Operations, infrastructure, inference, security, and data center teams
  • Understand platform users' needs and make practical product and engineering decisions
  • Lead projects from initial design to production, measure impact, and use operational feedback to guide improvements

Requirements

What you’ll need
  • 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
  • Strong Go or Python skills and experience designing services and APIs
  • Expertise in control planes, fleet management systems, or operational platforms
  • Experience with Linux, containers, Kubernetes, and handling failures across distributed systems
  • Experience designing systems that handle asynchronous work, retries, and partial failures
  • Experience with event streaming, workflow automation, or time-series telemetry
  • Strong judgment in reliability, security, and observability
  • Ability to lead ambiguous projects and work effectively across engineering and operations teams
  • Preferred: experience building software for incident response, hardware health, or capacity management
  • Preferred: experience developing dashboards and applications for infrastructure operators
  • Preferred: experience working with AI clusters and their compute, networking, and hardware systems

Benefits

Comp & perks
  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Simple, non-corporate work culture that respects individual beliefs
  • Continuous learning, growth and support
  • Equal and diverse work environment