Failure Analysis Engineer - Server Systems Integration
AMD
- Location
- Secaucus, New Jersey
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 2h ago
Skills
About this role
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.
THE ROLE
Own end-to-end FA across hardware, firmware, silicon, and integration for Helios server platform bring-up and system-level failures. This role drives firmware-aware hardware debug across BIOS, BMC, CPLD, PMBus, POST, PCIe, high-speed interconnect initialization, power sequencing, register state, and platform-level interactions. The engineer will determine whether failures are driven by firmware behavior, hardware design, silicon behavior, component quality, power delivery, configuration, or cross-domain interaction issues. Success in this role requires serving as the firmware SME for FA, enabling independent isolation of systemic system issues while owning interfaces with BIOS, BMC, silicon, validation, design, manufacturing, supplier quality, and customer-facing teams to drive corrective actions and improve platform quality. THE PERSON: The ideal candidate is a hands-on systems integrator and failure analysis technical leader with deep platform bring-up experience on complex server or hyperscale systems. They can navigate ambiguous failures across BIOS, BMC, CPLD, firmware, silicon, motherboard, PDB, power delivery, PCIe, high-speed interconnects, diagnostics, and system configuration boundaries. This person is not expected to be a firmware coder; instead, they must understand firmware-controlled hardware behavior well enough to isolate whether the failure is firmware, hardware, or an interaction between domains. They should be comfortable leading structured debug, reviewing register dumps and logs, validating behavior with scopes and logic analyzers, defining diagnostic strategy, communicating clear RCA conclusions, and driving corrective actions across cross-functional engineering teams. KEY RESPONSIBILITIES : Lead end-to-end failure analysis and root cause ownership for Helios server platform bring-up, factory, customer, and system-level failures. Own platform bring-up debug across BIOS, BMC, CPLD, PMBus, POST, PCIe enumeration, high-speed interconnect initialization, resets, clocks, power sequencing, and system configuration domains. Determine whether failures are caused by firmware behavior, hardware design, component quality, silicon behavior, power delivery, manufacturing process, configuration, or cross-domain interaction issues. Develop and execute structured debug plans using register dumps, firmware and BIOS logs, BMC event logs, telemetry, POST codes, diagnostic results, schematics, board layouts, oscilloscope captures, and logic analyzer traces. Perform power sequencing and platform readiness debug, including rail enable timing, reset behavior, clock availability, PMBus communication, voltage/current telemetry, and fault propagation analysis. Validate firmware-controlled hardware behavior using scopes, logic analyzers, protocol tools, register reads, and data-driven correlation across boot, initialization, and failure states. Own technical interfaces with silicon, BIOS, BMC, firmware, validation, diagnostics, hardware design, manufacturing, and supplier teams to lead cross-functional RCA and drive corrective action closure. Automate debug data collection, log parsing, register analysis, and failure correlation using Python, Linux tools, scripting, and data analysis workflows. Create clear technical reports, executive summaries, debug timelines, and 8D-style documentation that communicate failure mode, evidence, root cause, impact, and