Data Center Hardware Quality & Reliability Engineer

OpenAI · San Francisco · $226K – $285K · Engineering

Posted 2026-09-28

Apply for this role →

About The Role

Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.

The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.

Key Responsibilities

• Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.

• Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.

• Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.

• Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.

• Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.

• Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.

• Close the loop by verifying whether upstream changes reduce field recurrence.

• Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.

• Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.

• Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.

• Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers.

Qualifications

• BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred.

• 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes.

• Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required.

• Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth.

• Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification.

• Working proficiency with SQL and Python/R or equivalent analytics tools.

• Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority.

Preferred Skills

• GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations.

• Design for serviceability: FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy.

• Qualification-to-field correlation and mission-profile development.

• ODM/CM/supplier experience: FA quality, audit, QBR, and corrective-action governance.

• Linux/BMC/IPMI/Redfish logs and fleet telemetry.

• Leadership of a cross-generation reliability program or launch-readiness gate.

Apply for this role →

← Back to all jobs