FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production infrastructure operations, including monitoring, incident response, and automation. Proven ability to establish operational processes and lead technical teams in complex environments.
Highest-signal resume keywords
Infrastructure Operations ManagementMonitoring And TelemetryIncident Response CoordinationAutomation And Operational EfficiencyCross-Functional Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production Infrastructure ManagementOperational Process DevelopmentHardware TroubleshootingService RestorationFleet OperationsData Center OperationsIncident ManagementOperational AutomationRMA Process ManagementHPC Operations
Soft Skills
Cross-Functional CommunicationLeadershipOperational Judgment
Industry Keywords
Cloud OperationsAI/HPCLarge-Scale ComputeDistributed InfrastructureGeographically Distributed Infrastructure
Tech Stack
Tools & technologiesCloud
About the role
Key responsibilities & impact- Own the day-to-day operational health of Cerebras-managed production infrastructure
- Establish consistent operational processes across clusters, sites, and operating environments
- Maintain visibility into production state, degradation, outages, maintenance, and infrastructure availability
- Establish fleet monitoring, alert response, triage, escalation, and restoration processes
- Develop 24x7 operational coverage as fleet requirements evolve
- Establish issue ownership from detection through restoration
- Develop and mature the Cerebras Fleet Operations/NOC capability
- Establish standards for monitoring, telemetry, alerting, dashboards, and fleet-health reporting
- Develop shift, handoff, escalation, and on-call procedures
- Ensure operators have required tooling, runbooks, procedures, and access
- Partner with Service Ops & Enablement on operator competency and training
- Drive standardization across geographically distributed infrastructure
- Determine when physical intervention is required and dispatch/prioritize SiteOps/DC Ops activities
- Coordinate troubleshooting between centralized Operations and onsite technicians
- Validate recovery and authorize return-to-service
- Analyze recurring physical interventions for diagnostics, procedures, serviceability, and automation improvements
- Establish and manage failed-system and component repair/RMA workflows
- Coordinate diagnosis, replacement, repair, and RMA activities
- Manage failed-asset disposition and replacement-spares availability
- Track repair status and repair-loop performance
- Feed recurring repair patterns into Reliability & Incident Management and Engineering
- Establish governance for production maintenance and change
- Coordinate planned maintenance across Central Ops, SiteOps, customers, Engineering, and infrastructure partners
- Ensure production changes have validation, execution, and rollback procedures
- Maintain maintenance windows and communication mechanisms
- Identify and prioritize operational automation opportunities
- Provide requirements and acceptance criteria to Software/Engineering teams
- Increase automated detection, diagnosis, remediation, and validation
- Reduce unnecessary Engineering involvement and operational toil
Requirements
What you’ll need- 8+ years of experience in infrastructure operations, cloud operations, fleet operations, SRE, HPC operations, data center operations, or similar technical environments
- 3+ years leading technical operations teams
- Experience operating complex, highly available production infrastructure
- Strong understanding of monitoring, telemetry, incident response, maintenance, hardware troubleshooting, and service restoration
- Experience coordinating work between centralized operations and onsite technical teams
- Strong infrastructure troubleshooting and operational judgment
- Ability to establish repeatable processes in rapidly evolving environments
- Strong cross-functional leadership and communication skills
- Focus on automation and reduction of operational toil
- Preferred: AI/HPC, large-scale compute, accelerator, cloud, or distributed infrastructure experience
- Preferred: Experience with hardware-intensive production environments
- Preferred: Experience building or operating NOC/Fleet Operations capabilities
- Preferred: Experience managing hardware repair/RMA processes
- Preferred: Experience operating geographically distributed infrastructure
Benefits
Comp & perks- Build a breakthrough AI platform beyond the constraints of the GPU
- Publish and open source cutting-edge AI research
- Work on one of the fastest AI supercomputers in the world
- Job stability with startup vitality
- Simple, non-corporate work culture that respects individual beliefs
- Continuous learning, growth and support
- Equal employment opportunity environment
