FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production infrastructure operations, including monitoring, incident response, and automation. Proven ability to establish operational processes and lead cross-functional teams in complex environments.
Highest-signal resume keywords
Infrastructure OperationsCloud OperationsMonitoring and TelemetryIncident ResponseAutomation Focus
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production Infrastructure ManagementOperational Process DevelopmentHardware TroubleshootingService RestorationFleet OperationsRMA Process ManagementAutomation Opportunities IdentificationOperational JudgmentData Center OperationsHPC Operations
Soft Skills
Cross-Functional LeadershipCommunication SkillsTeam Coordination
Industry Keywords
AILarge-Scale ComputeDistributed InfrastructureNOC OperationsGeographically Distributed Infrastructure
Tech Stack
Tools & technologiesCloud
About the role
Key responsibilities & impact- Own day-to-day operational health of Cerebras-managed production infrastructure
- Establish consistent operational processes across clusters, sites, and operating environments
- Maintain visibility into production state, degradation, outages, maintenance, and infrastructure availability
- Establish fleet monitoring, alert response, triage, escalation, and restoration processes
- Develop 24x7 operational coverage as fleet requirements evolve
- Establish ownership of production issues from detection through restoration
- Develop and mature the Cerebras Fleet Operations/NOC capability
- Establish standards for monitoring, telemetry, alerting, dashboards, and fleet-health reporting
- Develop shift, handoff, escalation, and on-call procedures
- Ensure operators have required tooling, runbooks, procedures, and access
- Partner with Service Ops & Enablement on operator competency and training
- Drive standardization across geographically distributed infrastructure
- Determine when physical intervention is required and dispatch/prioritize SiteOps/DC Ops activities
- Coordinate troubleshooting between centralized Operations and onsite technicians
- Validate recovery and authorize return-to-service
- Analyze recurring physical interventions for diagnostics, procedures, serviceability, or automation improvements
- Establish and manage failed-system and component repair/RMA workflows
- Coordinate diagnosis, replacement, repair, and RMA activities
- Establish disposition paths for failed assets and maintain repair-loop visibility
- Coordinate planned maintenance and ensure production changes have validation, execution, and rollback procedures
- Identify and prioritize automation opportunities
- Provide operational requirements and acceptance criteria to Software/Engineering teams
- Increase automated detection, diagnosis, remediation, and validation
- Reduce unnecessary Engineering involvement in repeatable operational activities
- Report to the Director of Central Operations and collaborate with SiteOps/DC Ops, Reliability & Incident Management, Service Ops & Enablement, Global Service Logistics & Inventory, Build & Deploy, and Engineering
Requirements
What you’ll need- 8+ years of experience in infrastructure operations, cloud operations, fleet operations, SRE, HPC operations, data center operations, or similar technical environments
- 3+ years leading technical operations teams
- Experience operating complex, highly available production infrastructure
- Strong understanding of monitoring, telemetry, incident response, maintenance, hardware troubleshooting, and service restoration
- Experience coordinating work between centralized operations and onsite technical teams
- Strong infrastructure troubleshooting and operational judgment
- Demonstrated ability to establish repeatable processes in rapidly evolving environments
- Strong cross-functional leadership and communication skills
- Demonstrated focus on automation and reduction of operational toil
- Preferred: AI/HPC, large-scale compute, accelerator, cloud, or distributed infrastructure experience
- Preferred: Experience with hardware-intensive production environments
- Preferred: Experience building or operating NOC/Fleet Operations capabilities
- Preferred: Experience managing hardware repair/RMA processes
- Preferred: Experience operating geographically distributed infrastructure
Benefits
Comp & perks- Build a breakthrough AI platform beyond the constraints of the GPU
- Publish and open source cutting-edge AI research
- Work on one of the fastest AI supercomputers in the world
- Job stability with startup vitality
- Simple, non-corporate work culture that respects individual beliefs
- Equal employment opportunity environment
- Continuous learning, growth and support
