FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Center Hardware Quality – Reliability Engineer
OpenAI. Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-generation platforms .
Posted 9/28/2026full-timeSan Francisco • California • United StatesSeniorLead💰 $226,000 - $285,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in hardware quality and reliability engineering, with a strong focus on field-failure analysis, reliability statistics, and data-driven decision-making. Proficient in leading cross-functional teams and mentoring engineers to enhance product reliability and serviceability.
Highest-signal resume keywords
Hardware Quality/Reliability EngineeringField-Failure AnalysisReliability StatisticsFMEA/FTASQL and Python/R Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reliability StatisticsFMEAFTAField-Failure AnalysisCorrective-Action VerificationReliability GrowthTelemetry AnalysisData ModelingFailure AnalysisMetrics Definition
Soft Skills
Influencing Without AuthorityMentoringCross-Functional Collaboration
Tools & Technologies
SQLPythonLinuxBMCIPMIRedfishTelemetry ToolsAnalytics Tools
Industry Keywords
Data-Center OperationsMission-Critical InfrastructureRMACAPAHigh-Power DeliveryGPU/AI Server PlatformsDesign for ServiceabilitySupplier ExperienceAuditQBR
Tech Stack
Tools & technologiesLinuxPythonSQL
About the role
Key responsibilities & impact- Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-generation platforms
- Build and govern the field-quality data model across telemetry, tickets, RMA/repair, failure analysis, firmware, configuration, supplier, and manufacturing genealogy
- Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty
- Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations
- Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring
- Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age
- Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates
- Verify whether upstream changes reduce field recurrence
- Define supplier/CM failure-analysis standards, field-data contracts, scorecards, escalation paths, and closure evidence
- Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement
- Create executive decision packages covering population at risk, exposure, confidence, options, cost/risk, and recommendation
- Run the cross-functional reliability council and mentor 1P and 3P Field Quality Engineers as the team grows
Requirements
What you’ll need- BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred
- 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure
- 3+ years owning field-failure, RMA, or CAPA outcomes
- Working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior
- Reliability statistics knowledge: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth
- Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, failure analysis, and corrective-action verification
- Working proficiency with SQL and Python/R or equivalent analytics tools
- Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority
- Preferred: GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations
- Preferred: design for serviceability, FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy
- Preferred: qualification-to-field correlation and mission-profile development
- Preferred: ODM/CM/supplier experience in FA quality, audit, QBR, and corrective-action governance
- Preferred: Linux/BMC/IPMI/Redfish logs and fleet telemetry
- Preferred: leadership of a cross-generation reliability program or launch-readiness gate
Benefits
Comp & perks- Equity
- Performance-related bonus(es) for eligible employees
- Medical, dental, and vision insurance for you and your family
- Employer contributions to Health Savings Accounts
- Pre-tax Health FSA, Dependent Care FSA, and commuter expense accounts
- 401(k) retirement plan with employer match
- Paid parental leave up to 24 weeks for birth parents and 20 weeks for non-birthing parents
- Paid medical and caregiver leave up to 8 weeks
- Flexible PTO for exempt employees
- Up to 15 days annually for non-exempt employees
- 13+ paid company holidays
- Paid coordinated company office closures
- Paid sick or safe time
- Mental health and wellness support
- Employer-paid basic life and disability coverage
- Annual learning and development stipend
- Daily meals in offices
- Meal delivery credits as eligible
- Relocation support for eligible employees
- Charitable donation matching and wellness stipends may be provided