Skip to main content

OpenAI

Data Center Hardware Quality & Reliability Engineer

San Francisco (Remote)$226k–$285kfulltimemidAdded today

About this role

Lead hardware quality and reliability for OpenAI's data-center infrastructure, managing the complete lifecycle from field failures to root-cause analysis and preventive design changes. Combine hardware expertise, statistical reliability methods, and cross-functional leadership to reduce fleet failures, improve manufacturing quality, and forecast spares across both third-party and internal platforms.

What you'll do

  • Build and maintain field-quality data models integrating telemetry, RMA, failure analysis, firmware, and supplier data
  • Define and track reliability metrics (AFR, ASR, MTBF/MTTR, DPPM) with clear denominators and uncertainty quantification
  • Lead end-to-end triage, containment, root-cause analysis, 8D/CAPA, and corrective-action verification for field failures
  • Develop spare-demand and reliability-growth projections by product, component, supplier, and geography
  • Partner with manufacturing and design teams to convert field failure modes into improved test coverage and design requirements
  • Establish supplier standards, scorecards, and escalation governance; mentor growing Field Quality team

What they're looking for

  • Hardware/system architecture (board, rack, firmware, manufacturing test)
  • Reliability statistics (Weibull, censored data, confidence bounds, MTBF/MTTR)
  • FMEA/FTA, 8D/CAPA, and failure analysis methodologies
  • SQL and Python/R for data analysis and modeling
  • Cross-functional influence and technical leadership without direct authority
  • Design for serviceability and repair-workflow optimization
  • Fleet telemetry and BMC/IPMI/Redfish log analysis
  • Server/mission-critical infrastructure operations
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a significant field failure you owned: how did you triage it, establish root cause, and verify the corrective action prevented recurrence?
  • How would you build a field-quality data model that integrates telemetry, RMA tickets, and supplier genealogy for a fleet with multiple hardware generations?