yoinka

System Hardware Reliability Engineer

Google

Sunnyvale, CA, USASenior$188k – $274k/yrH-1B sponsor company
Sign in to applyVerified 1h ago
Location
Sunnyvale, CA, USA
Work model
On-Site
Level
Senior
Salary
$188k – $274k/yr
H-1B history
2,460 approvals (FY2023)
Posted
1h ago

Skills

GCPMachine LearningPyTorchTensorFlow

About this role

As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations. As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems. You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint. The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide. We're the driving team behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $188000 - $274000 (USD) + 20% bonus target + equity + benefits Learn more about benefits at Google .

Design and implement Prognostics and Health Management (PHM) algorithms using physics-informed machine learning to forecast hardware degradation and Remaining Useful Life (RUL). Develop stochastic degradation models to predict the impact of dynamic thermal and power envelopes on fleet reliability. Build health state monitoring and anomaly detection frameworks leveraging massive fleet telemetry data to enable predictive maintenance. Partner with software and controls teams to integrate predictive health models into automated load-management and thermal capping mechanisms. Lead comprehensive reliability assessments and PoF modeling for silicon, interconnects, optics, thermal solutions and power delivery/battery systems. Drive Failure Modes and Effects Analyses (FMEAs) to identify vulnerabilities under extreme environmental operating conditions.

Minimum qualifications: Bachelor’s degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent practical experience. 8 years of experience in applying Design for Reliability techniques, and working on multiple consumer electronics products. 8 years of experience in hardware reliability engineering, physics of failure, and predictive analytics. Preferred qualifications:

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

System Hardware Reliability Engineer at Google, Sunnyvale, CA, USA | Yoinka