Hardware Reliability Engineer
Description
As a member of Meta Infrastructure's Hardware Product Integrity team you will work on next-generation data center hardware. You will be a part of futuristic projects including HW that will serve as the backbone for Meta's AGI vision, be in a position to influence HW technology that serves to connect billions of people across the world! In this role, you will drive reliability engineering efforts across server, storage, and networking hardware deployed in Meta's data centers, applying failure analysis, accelerated life testing, and reliability modeling to reduce field failures and improve hardware reliability.
Responsibilities
Lead DFR activities such as DFMEA, derating across various AI, compute and storage platforms Understanding technology that drives compute, storage, server hardware, and networking modules to develop reliability tests to bring out design weaknesses Establish Design Verification tests, to bring out environmental stress weaknesses in server design and ensure designs meet Meta's lifetime reliability metrics Work closely with ODMs to ensure and oversee tests are being executed as planned, suggest necessary improvements based on lessons learned from previous platforms Translate test results into meaningful product life metrics and highlight shortcomings in any metrics that are not met Utilize reliability statistics to help with decision making and quantifying risk and Lead the development of internal reliability test infrastructure to support initiatives and design of experiments Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues
Qualifications
Bachelor's degree in Electrical Engineering or Mechanical Engineering or a related discipline 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations Experience in silicon reliability and working on custom silicon is a plus MSc in Mechanical or Electrical Engineering or related disciplines Familiarity with data center environments is beneficial First-hand knowledge of server rack hardware is preferred
Compensation: $144,000/year to $204,000/year + bonus + equity + benefits