Hardware Reliability Engineer

GoogleApplyPublished 1 days agoFirst seen 1 days ago
Apply
As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform design of new products. Your strong interpersonal and communication skills will be key to ensuring adoption of your technical recommendations.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $159000 - $230000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Lead system design evaluations and implement reliability plans to assess and mitigate failure risks early in New Product Introduction (NPI).
  • Drive reliability test plans and collect, analyze, and synthesize test data to enable verification of design reliability goals.
  • Lead system reliability monitoring efforts (availability, repair trends) and proactively alert product teams on unwanted system behavior, working on mitigation strategy definition and implementation.
  • Extract field reliability data to drive root-cause failure analysis while managing global Contract Manufacturer (CM) reliability and Ongoing Reliability Test (ORT) execution.
  • Manage efforts with outside partners, testing labs, failure analysis labs, and cross-functional internal groups, while developing in-house test and qualification capabilities where needed.

Minimum qualifications:

  • Bachelor's degree in Hardware Engineering, or equivalent practical experience.
  • 6 years of experience in hardware reliability engineering, physics of failure, and predictive analytics.
  • 5 years of experience in reliability engineering of cloud infrastructure hardware and technology, failure analysis, and fault isolation techniques.
  • 3 years of experience in rigid PCB manufacturing processes.
  • Experience in statistics and data analysis using one or more software/tools (e.g., MATLAB, Python, JMP, or Minitab).

Preferred qualifications:

  • Master's degree or PhD in Hardware Engineering.
  • Experience with system level reliability tools such as Reliability Block Diagrams (RBDs), Mean Cumulative Function (MCF), Homogeneous and Non-Homogeneous Poisson Processes (HPP, NHPP), and simulation tools.
  • Experience leading cross-functional, problem-solving teams using practical approaches.
  • Ability to effectively lead teams to meet corporate and customer reliability expectations.
  • Ability to effectively communicate to the project team with excellent people management skills.
  • Excellent leadership, de-escalation, and executive communication skills.