Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Staff Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC

. Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues .

Posted 10/2/2026full-timeBellevue • Washington • United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in monitoring and maintaining AI infrastructure, with a strong focus on incident response, troubleshooting, and operational excellence. Proficient in collaborating with cross-functional teams to enhance operational tooling and processes in a dynamic environment.

Highest-signal resume keywords
Data Center Operations ExperienceIncident Response SkillsGPU Infrastructure SupportMonitoring Stack ProficiencyOn-Call Management Tooling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Incident ResponseTroubleshootingData Center OperationsSite Reliability EngineeringGPU SupportOperational ToolingMonitoringCapacity ManagementPerformance MonitoringRoot Cause Analysis
Soft Skills
CollaborationProblem-SolvingAdaptabilityCommunicationCustomer Support
Tools & Technologies
PagerDutyOpsgeniePrometheusGrafanaDatadog
Industry Keywords
AI InfrastructureOperational MaturityIncident ManagementLarge-Scale ComputeEarly-Stage Environment

Tech Stack

Tools & technologies
GrafanaPrometheus

About the role

Key responsibilities & impact
  • Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues
  • Lead or support incident response for production issues and drive fast, effective resolution
  • Collaborate with engineering teams to ensure monitoring, alerting, and operational tooling are first class and identify issues before customer impact
  • Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence
  • Contribute to runbooks, on-call practices, and operational maturity as the platform scales
  • Help build and scale the Operations & Maintenance function supporting mission-critical AI infrastructure
  • Support post-launch monitoring, incident response, reliability, and exceptional customer support
  • Build operations and maintenance processes, tooling, systems, and operational discipline for next-generation AI infrastructure

Requirements

What you’ll need
  • Experience in a data center operations, site reliability, or infrastructure operations role
  • Experience ideally supporting GPU or large-scale compute environments
  • Strong incident response and troubleshooting skills across hardware, networking, and systems layers
  • Comfortable working in an early-stage environment where processes and tooling are still being established
  • Experience in multiple technical environments is a plus
  • Experience with on-call/incident-management tooling such as PagerDuty or Opsgenie is preferred
  • Experience with monitoring stacks such as Prometheus, Grafana, or Datadog is preferred
  • Background supporting GPU cluster or data center operations post-launch is preferred
  • U.S. work authorization is required
  • Visa sponsorship is not currently available
  • Willingness to work at least three days per week in-office in downtown Bellevue, WA

Benefits

Comp & perks
  • Certain roles are eligible for merit increases, annual bonus, and long-term incentives based on individual performance
  • Medical, dental, and vision insurance for U.S.-based employees
  • 401(k) plan and company match
  • Paid holidays per calendar year
  • Hybrid work arrangement