Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC

. Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues .

Posted 10/2/2026full-timeBellevue • Washington • United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in monitoring and incident response for data center operations, particularly in GPU and large-scale compute environments. Proficient in troubleshooting across hardware, networking, and systems layers while contributing to operational maturity and customer support.

Highest-signal resume keywords
Data Center OperationsIncident ResponseTroubleshooting SkillsMonitoring StacksOn-Call Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Incident ManagementTroubleshootingPerformance MonitoringCapacity ManagementReliability Engineering
Soft Skills
CollaborationProblem-SolvingAdaptability
Tools & Technologies
PagerDutyOpsgeniePrometheusGrafanaDatadog
Industry Keywords
Site ReliabilityInfrastructure OperationsGPU InfrastructureOperational MaturityMission-Critical Systems

Tech Stack

Tools & technologies
GrafanaPrometheus

About the role

Key responsibilities & impact
  • Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues
  • Lead or support incident response for production issues and drive fast, effective resolution
  • Collaborate with engineering teams to ensure monitoring, alerting, and operational tooling are first class and catch issues before customer impact
  • Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence
  • Contribute to runbooks, on-call practices, and operational maturity as the platform scales
  • Help build and scale the Operations & Maintenance function supporting mission-critical AI infrastructure
  • Support post-launch monitoring, incident response, reliability, and exceptional customer support

Requirements

What you’ll need
  • Experience in a data center operations, site reliability, or infrastructure operations role, ideally supporting GPU or large-scale compute environments
  • Strong incident response and troubleshooting skills across hardware, networking, and systems layers
  • Comfortable working in an early-stage environment where processes and tooling are still being established
  • Experience in multiple technical environments is a plus
  • Experience with on-call/incident-management tooling such as PagerDuty or Opsgenie is preferred
  • Experience with monitoring stacks such as Prometheus, Grafana, or Datadog is preferred
  • Background supporting GPU cluster or data center operations post-launch is preferred
  • U.S. work authorization is required
  • Visa sponsorship is not currently available
  • Minimum of three days per week in-office once the permanent office is established

Benefits

Comp & perks
  • Certain roles are eligible for merit increases, annual bonus, and long term incentives, allocated based on individual performance
  • Medical, dental, and vision insurance for U.S.-based employees
  • 401(k) plan and company match
  • Paid holidays per calendar year
  • Hybrid work arrangement