OpenAI · Infrastructure · Staff+ · Posted 2026-09-28
Data Center Hardware Quality & Reliability Engineer
OpenAI · San Francisco · $226k–285k base
This range's midpoint is above 73% of posted infrastructure ranges at AI companies right now. See the salary index.
Apply on OpenAI's site Watch OpenAI for new roles
About The Role
Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.
The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.
Key Responsibilities
• Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.
• Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.
• Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.
• Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.
• Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.
• Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.
• Close the loop by verifying whether upstream changes reduce field recurrence.
• Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.
• Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.
• Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.
• Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers.
Qualifications
• BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred.
• 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes.
• Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required.
• Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth.
• Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification.
• Working proficiency with SQL and Python/R or equivalent analytics tools.
• Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority.
Preferred Skills
• GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations.
• Design for serviceability: FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy.
• Qualification-to-field correlation and mission-profile development.
• ODM/CM/supplier experience: FA quality, audit, QBR, and corrective-action governance.
• Linux/BMC/IPMI/Redfish logs and fleet telemetry.
• Leadership of a cross-generation reliability program or launch-readiness gate.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the ma …
More infrastructure roles at OpenAI
-
Platform Engineering Manager, Forward Deployed Engineering (FDE)
InfrastructureLead / ManagerRemote US$302k–335ktoday
-
Software Engineer, DevOps
InfrastructureRemote US$177k–327ktoday
-
Technical Program Manager, Hardware Systems
InfrastructureUS$207k–242ktoday
-
Software Engineer, Search Infrastructure
InfrastructureRemote US$266k–445k3d
-
Rack Power Engineer
InfrastructureStaff+Remote US$287k–485k5d
See also: AI jobs in San Francisco Bay Area · OpenAI salaries · Python jobs · SQL jobs · Speech jobs.
This listing is reproduced from OpenAI's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All OpenAI roles · AI salaries.