yoinka

Senior Cloud Hardware Storage Engineer

Microsoft

United States, California, Aliso Viejo; United States, Washington, RedmondSeniorH-1B sponsor company
Sign in to applyVerified 1h ago
Location
United States, California, Aliso Viejo; United States, Washington, Redmond
Work model
On-Site
Level
Senior
H-1B history
2,066 approvals (FY2023)
Posted
2h ago

Skills

Azure

About this role

Overview

Microsoft Silicon and Cloud Hardware Infrastructure Engineering (SCHIE) is the team behind Microsoft’s expanding Cloud Infrastructure and responsible for powering Microsoft’s “Intelligent Cloud” mission. CHIE delivers the core infrastructure and foundational technologies for Microsoft's over 200 online businesses including Bing, MSN, Office 365, Xbox Live, Skype, OneDrive and the Microsoft Azure platform globally with our server and data center infrastructure, security and compliance, operations, globalization, and manageability solutions. Our focus is on smart growth, high efficiency, and delivering a trusted experience to customers and partners worldwide and we are looking for passionate, high-energy engineers to help achieve that mission.         As Microsoft's cloud business continues to grow the ability to deploy new offerings and HW infrastructure on time, in high volume with high quality and lowest cost is of paramount importance. To achieve this goal, the Silicon Cloud Hardware Infrastructure Engineering (SCHIE) team is instrumental in defining and delivering measures of success for hardware design, qualification, fleet support, scale, and sustainability related to Microsoft cloud hardware.       Azure Memory and Storage Center of Excellence (AMS CoE) is part of the SCHIE organization focusing on Memory and Storage devices going into the Cloud hardware servers. AMS provide memory and storage solutions to Azure, drive memory and storage suppliers to deliver high quality products, meeting our requirements.        We are looking for a Senior Cloud Hardware Engineer to scale Azure’s Fault Self‑Healing and Failure Prediction systems.     You will own the end‑to‑end technical design and execution of the fault prevention ecosystem, spanning telemetry, automation, isolation logic, firmware interactions, and repair workflows, operating at hyperscale across millions of nodes. The role directly impacts customer uptime and fleet availability.

Responsibilities

Design and build best-in-class fleet resiliency systems for storage devices at scale     Develop scalable live monitoring capabilities, fault detection and repair solutions   Design features for SSDs, HDDs and Storage Accelerator firmware deployment   Lead collaboration projects with hardware, firmware and software teams that fault reduction projects  Build automation to drive repair efficiency for storage operations in the production fleet  Collaborate with suppliers to design reliable, high performance and quality storage devices     Analyze data to identify, prototype, and drive the implementation of technical and process improvements to increase the predictability, agility, and quality of Azure systems     Actively support Azure service stakeholders   Qualifications Required/minimum qualifications Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience   Other Requirements: Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications

B.S. Degree in Computer Engineering, Computer Science, Electrical or equivalent experience   10+ years of SSD firmware engineering development experience  6+ years of NVMe and PCIe experience  Deep expertise in storage device resiliency, fault analysis, and live‑site operations.      Lead end‑to‑end design decisions across detection, prediction,

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Cloud Hardware Storage Engineer at Microsoft, United States, California, Aliso Viejo; United States, Washington, Redmond | Yoinka