FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Principal Data Center Operations and Maintenance Engineer
Designworks Talent LLC. Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in monitoring and maintaining AI infrastructure, with a strong focus on incident response, troubleshooting, and operational excellence. Proficient in collaborating with cross-functional teams to enhance operational tooling and processes in a dynamic environment.
Highest-signal resume keywords
Data Center Operations ExperienceIncident Response SkillsGPU Infrastructure SupportMonitoring Stack ProficiencyOn-Call Management Tooling
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Incident ResponseTroubleshootingData Center OperationsSite Reliability EngineeringGPU SupportOperational ToolingMonitoringCapacity ManagementPerformance MonitoringRoot Cause Analysis
Soft Skills
CollaborationProblem-SolvingAdaptabilityCommunicationCustomer Support
Tools & Technologies
PagerDutyOpsgeniePrometheusGrafanaDatadog
Industry Keywords
AI InfrastructureOperational MaturityIncident ManagementLarge-Scale ComputeEarly-Stage Environment
Tech Stack
Tools & technologiesGrafanaPrometheus
About the role
Key responsibilities & impact- Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues
- Lead or support incident response for production issues and drive fast, effective resolution
- Collaborate with engineering teams to ensure monitoring, alerting, and operational tooling are first class and identify issues before customer impact
- Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence
- Contribute to runbooks, on-call practices, and operational maturity as the platform scales
- Help build and scale the Operations & Maintenance function supporting mission-critical AI infrastructure
- Support post-launch monitoring, incident response, reliability, and exceptional customer support
- Build operations and maintenance processes, tooling, systems, and operational discipline for next-generation AI infrastructure
Requirements
What you’ll need- Experience in a data center operations, site reliability, or infrastructure operations role
- Experience ideally supporting GPU or large-scale compute environments
- Strong incident response and troubleshooting skills across hardware, networking, and systems layers
- Comfortable working in an early-stage environment where processes and tooling are still being established
- Experience in multiple technical environments is a plus
- Experience with on-call/incident-management tooling such as PagerDuty or Opsgenie is preferred
- Experience with monitoring stacks such as Prometheus, Grafana, or Datadog is preferred
- Background supporting GPU cluster or data center operations post-launch is preferred
- U.S. work authorization is required
- Visa sponsorship is not currently available
- Willingness to work at least three days per week in-office in downtown Bellevue, WA
Benefits
Comp & perks- Certain roles are eligible for merit increases, annual bonus, and long-term incentives based on individual performance
- Medical, dental, and vision insurance for U.S.-based employees
- 401(k) plan and company match
- Paid holidays per calendar year
- Hybrid work arrangement