Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering
Amazon
- Location
- US, CA, Cupertino
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- Aug 27, 2026
Skills
About this role
Application deadline: Sep 1, 2026 Do you want to build the infrastructure that keeps Artificial Intelligence compute capacity available to Generative AI customers? Do you want to solve problems at the boundary between physical hardware and software - at cloud scale? AWS Hardware Engineering is looking for a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale. You will develop automation, analyze hardware telemetry across tens of thousands of hosts, and build tooling that directly determines whether capacity is available to customers. Your work spans the full stack — from hardware monitoring interfaces, to health diagnostics up through fleet-wide data pipelines and operational dashboards. Key job responsibilities Fleet Health & Data Analysis - Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact - Contribute to predictive failure detection using sensor data, error trending, and log correlation - Build and maintain operational dashboards and metrics for platform fleet health. - Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets Systems Development & Automation - Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows - Develop diagnostic tools for Linux on ARM and x86 architectures - Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems - Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability Cross-Team Collaboration - Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions - Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers - Participate in sprint-based planning and oncall rotation for platform-level escalations A day in the life Some days you are deep in system event logs chasing a failure pattern across thousands of hosts; other days you are writing automation that eliminates a manual triage workflow entirely. You work with hardware engineers, firmware teams, datacenter operations, and vendor partners - driving quality and reliability from manufacturing through steady-state operations. Located in Cupertino, Seattle, or Denver, you work with global development teams on servers deployed in datacenters worldwide.
About the team
AWS Hardware Engineering designs and delivers next-generation cloud infrastructure - the servers, accelerators, and storage platforms that power AWS. Our team builds custom systems for AI training, inference, and compute workloads at global scale. We are directly responsible for launching and maintaining server hardware in the fleet, working across internal development teams, and design partners. We value work-life harmony, inclusive culture, and continuous learning. Even if you do not meet all preferred qualifications listed below, we encourage you to apply.