FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Director, Data Center Operations and Maintenance
Designworks Talent LLC. Build, lead, mentor, and grow high-performing Operations & Maintenance teams .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and leading high-performing Operations and Maintenance teams, with a strong focus on operational strategy, incident management, and service reliability within large-scale AI infrastructure. Proven ability to drive operational excellence through automation, monitoring, and continuous improvement while effectively communicating with executive leadership.
Highest-signal resume keywords
Operations LeadershipIncident ManagementService ReliabilityCloud Infrastructure ExperienceExecutive Communication Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Operational Strategy DevelopmentIncident Response ProcessesChange ManagementProblem ManagementService Level Objectives DefinitionOperational KPIsRoot Cause AnalysisOperational GovernanceScalable Support ModelsOperational Maturity Frameworks
Soft Skills
MentoringTeam BuildingInfluencing LeadershipCross-Functional CollaborationAdaptability in Early-Stage Environments
Tools & Technologies
GrafanaPrometheusDatadogPagerDutyOpsgenie
Industry Keywords
AI InfrastructureHyperscale EnvironmentsProduction EngineeringData Center OperationsDistributed Systems
Tech Stack
Tools & technologiesCloudDistributed SystemsGrafanaPrometheus
About the role
Key responsibilities & impact- Build, lead, mentor, and grow high-performing Operations & Maintenance teams
- Develop the operational strategy, organizational structure, and execution model for large-scale AI infrastructure
- Establish operational processes for incident response, change management, problem management, and service reliability
- Lead major incident management efforts and executive communications during production events
- Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement
- Partner with Engineering on operational readiness for infrastructure deployments and platform launches
- Define service level objectives, operational KPIs, and reliability metrics across the infrastructure portfolio
- Build scalable on-call programs, escalation models, runbooks, and operational governance
- Champion root cause analysis and long-term corrective actions to improve platform resilience
- Influence infrastructure architecture and operational tooling to improve availability, efficiency, and customer experience
- Help shape the long-term operations organization as the company expands globally
- Partner with Engineering, Infrastructure, Networking, Hardware, and Customer Operations leadership
Requirements
What you’ll need- Experience leading Operations, Site Reliability, Infrastructure Operations, Data Center Operations, or Production Engineering organizations
- Proven success building or scaling operations teams within cloud infrastructure, hyperscale environments, AI infrastructure, or large distributed systems
- Deep expertise in production operations, incident management, service reliability, and operational excellence
- Experience leading cross-functional teams during high-severity production incidents
- Strong understanding of infrastructure operations across compute, networking, storage, and hardware environments
- Demonstrated success building operational processes, organizational structure, and scalable support models in high-growth environments
- Executive-level communication skills with the ability to influence engineering and business leadership
- Comfortable operating in an early-stage organization where many systems and processes are being built for the first time
- U.S. work authorization required
- Visa sponsorship is not currently available
- Approximately three days per week in the office
- Preferred: experience supporting hyperscale cloud platforms, GPU infrastructure, AI platforms, HPC environments, or large-scale data centers
- Preferred: experience with Grafana, Prometheus, Datadog, or similar observability platforms
- Preferred: familiarity with PagerDuty, Opsgenie, or equivalent incident management platforms
- Preferred: experience implementing operational maturity frameworks and reliability engineering best practices
- Preferred: track record of building globally distributed operations organizations
Benefits
Comp & perks- Certain roles are eligible for merit increases
- Annual bonus eligibility for certain roles
- Long term incentives for certain roles
- Medical insurance for U.S.-based employees
- Dental insurance for U.S.-based employees
- Vision insurance for U.S.-based employees
- 401(k) plan and company match
- Paid holidays per calendar year
- Hybrid work arrangement