Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cerebras

Senior Manager, Production & Fleet Operations

Cerebras

. Own day-to-day operational health of Cerebras-managed production infrastructure .

Posted 10/9/2026full-timeSunnyvale • California • United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing production infrastructure operations, including monitoring, incident response, and automation. Proven ability to establish operational processes and lead cross-functional teams in complex environments.

Highest-signal resume keywords
Infrastructure OperationsCloud OperationsMonitoring and TelemetryIncident ResponseAutomation Focus

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production Infrastructure ManagementOperational Process DevelopmentHardware TroubleshootingService RestorationFleet OperationsRMA Process ManagementAutomation Opportunities IdentificationOperational JudgmentData Center OperationsHPC Operations
Soft Skills
Cross-Functional LeadershipCommunication SkillsTeam Coordination
Industry Keywords
AILarge-Scale ComputeDistributed InfrastructureNOC OperationsGeographically Distributed Infrastructure

Tech Stack

Tools & technologies
Cloud

About the role

Key responsibilities & impact
  • Own day-to-day operational health of Cerebras-managed production infrastructure
  • Establish consistent operational processes across clusters, sites, and operating environments
  • Maintain visibility into production state, degradation, outages, maintenance, and infrastructure availability
  • Establish fleet monitoring, alert response, triage, escalation, and restoration processes
  • Develop 24x7 operational coverage as fleet requirements evolve
  • Establish ownership of production issues from detection through restoration
  • Develop and mature the Cerebras Fleet Operations/NOC capability
  • Establish standards for monitoring, telemetry, alerting, dashboards, and fleet-health reporting
  • Develop shift, handoff, escalation, and on-call procedures
  • Ensure operators have required tooling, runbooks, procedures, and access
  • Partner with Service Ops & Enablement on operator competency and training
  • Drive standardization across geographically distributed infrastructure
  • Determine when physical intervention is required and dispatch/prioritize SiteOps/DC Ops activities
  • Coordinate troubleshooting between centralized Operations and onsite technicians
  • Validate recovery and authorize return-to-service
  • Analyze recurring physical interventions for diagnostics, procedures, serviceability, or automation improvements
  • Establish and manage failed-system and component repair/RMA workflows
  • Coordinate diagnosis, replacement, repair, and RMA activities
  • Establish disposition paths for failed assets and maintain repair-loop visibility
  • Coordinate planned maintenance and ensure production changes have validation, execution, and rollback procedures
  • Identify and prioritize automation opportunities
  • Provide operational requirements and acceptance criteria to Software/Engineering teams
  • Increase automated detection, diagnosis, remediation, and validation
  • Reduce unnecessary Engineering involvement in repeatable operational activities
  • Report to the Director of Central Operations and collaborate with SiteOps/DC Ops, Reliability & Incident Management, Service Ops & Enablement, Global Service Logistics & Inventory, Build & Deploy, and Engineering

Requirements

What you’ll need
  • 8+ years of experience in infrastructure operations, cloud operations, fleet operations, SRE, HPC operations, data center operations, or similar technical environments
  • 3+ years leading technical operations teams
  • Experience operating complex, highly available production infrastructure
  • Strong understanding of monitoring, telemetry, incident response, maintenance, hardware troubleshooting, and service restoration
  • Experience coordinating work between centralized operations and onsite technical teams
  • Strong infrastructure troubleshooting and operational judgment
  • Demonstrated ability to establish repeatable processes in rapidly evolving environments
  • Strong cross-functional leadership and communication skills
  • Demonstrated focus on automation and reduction of operational toil
  • Preferred: AI/HPC, large-scale compute, accelerator, cloud, or distributed infrastructure experience
  • Preferred: Experience with hardware-intensive production environments
  • Preferred: Experience building or operating NOC/Fleet Operations capabilities
  • Preferred: Experience managing hardware repair/RMA processes
  • Preferred: Experience operating geographically distributed infrastructure

Benefits

Comp & perks
  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Simple, non-corporate work culture that respects individual beliefs
  • Equal employment opportunity environment
  • Continuous learning, growth and support