Principal AI Software Engineer
Microsoft (Eightfold Apply)
- Location
- United States, Washington, Redmond; United States, California, Mountain View; United States, Oregon, Hillsboro
- Work model
- On-Site
- Level
- Principal
- Posted
- 2h ago
Skills
About this role
Overview
Do you want to be at the forefront of innovating the latest hardware and systems designs to propel Microsoft’s cloud growth? Are you seeking a unique career opportunity that combines technical capabilities, cross team collaboration, with business insight and strategy? The SPARC organization is responsible for strategy, planning and architecture pathfinding, and manages Azure’s hardware roadmap from architecture concept, through production for Microsoft’s current and future offerings. The CSA team within SPARC is at the forefront of systems architecture and technology pathfinding spanning compute, memory, storage, and system interconnects. Drawing on deep insights of workloads, emerging technology trends, and focused industry engagements, CSA team’s charter is to define and evaluate novel systems architecture innovations through hardware/software co-design and advance them through technical readiness for productization. Join our Compute System Architecture (CSA) team within the System Planning and Architecture (SPARC) organization in Azure Hardware Systems & Infrastructure (AHSI). AHSI is the team behind Microsoft’s expanding cloud business, responsible for delivering the hardware systems and infrastructure for cloud computing across Microsoft Azure, Bing, MSN, Office 365, OneDrive, Skype, Teams and Xbox Live. The CSA team is seeking a Principal AI Software Engineer !
Responsibilities
Lead full system software prototyping to develop capable proof-of-concepts to evaluate hardware/software co-designed capabilities for memory TCO reduction such as through memory tiering/pooling and overcommit solutions for Azure usages and deployment scenarios. Develop deep insights through workload characterization and correlation to identify systems optimization opportunities. Collaborate with diverse workload experts across Microsoft and partner ISVs to engineer TCO-optimized solutions for Azure general-purpose and specialized compute fleet. Influence and shape hardware architecture and industry alignment, targeting three-to-six-year timeframe, with data-driven analysis, insights and recommendations. Lead characterization and optimization of Large Language Model (LLM) inference workloads with focus on KV Cache capacity, placement, migration, and utilization across GPU HBM, host DRAM, CXL memory expansion/pooling, SSD, and emerging memory tiers. Develop proof-of-concepts and evaluation frameworks to assess memory-tiering architectures for AI inference, including CXL pooled memory, memory expansion solutions, context-memory platforms, SSD-backed cache tiers, and hardware/software co-designed approaches for reducing inference TCO. Design and execute workload characterization studies for agentic, multi-turn, coding, reasoning, and long-context AI workloads to quantify memory consumption, latency, throughput, token efficiency, and system utilization. Analyze end-to-end data movement across GPU, CPU, storage, and networking subsystems, identifying optimization opportunities within GPU Direct Storage (GDS), GPU Direct RDMA (GDR), peer-to-peer memory transfers, and distributed inference pipelines. Develop software prototypes, framework extensions, and instrumentation to evaluate KV Cache offload, prefetching, migration, compression, deduplication, and memory-overcommit techniques. Build performance models and simulation frameworks to predict the impact of memory hierarchy innovations on large-scale inference deployments.
Qualifications
Required/minimum qualifications Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Other Qualifications: Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security