Manager, Solutions Architecture - AI Labs
NVIDIA
- Location
- US CA Santa Clara
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 394 approvals (FY2023)
- Posted
- 19h ago
Skills
About this role
NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at one of NVIDIA's most strategic AI lab customers, while directly contributing to solution design and deep-dive debugging and product roadmap shaping through customer feedback.
What you will be doing
Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments for frontier AI labs. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts. Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for a strategic AI lab. Personally lead and guide complex compute and network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters. Guide your team as they lead compute, network, and software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed. Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU and Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions. Partner with customer engineering and security stakeholders to understand security requirements and translate them into scalable AI infrastructure What we need to see: BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other Engineering fields or equivalent experience. 8+ overall years in Systems/Solutions/Field Engineering, Network or Data Center Engineering, or similar roles, with 2+ years leading or mentoring engineers or architects (formal manager or strong tech lead). Direct people management & recruiting experience for geographically distributed technical teams. System-level expertise across CPU/GPU server architecture, NICs, Linux, system software, and kernel drivers. Experience with data center networking including Ethernet and/or InfiniBand switches, NICs, fabrics, associated tooling, and cluster performance troubleshooting. Familiarity with data center infrastructure (power, cooling, deployment constraints). Proven ability to lead technical teams, set priorities, and drive complex projects from design through production. Demonstrated success working with Product Management, Sales, and Engineering. Strong time management skills and ability to balance planning with hands-on support where needed. Excellent written and verbal communication, including the ability to lead customer meetings, communicate status and risks, and produce clear design docs, debug summaries, and presentations. Ways to stand out from the crowd: Track record leading bring-up and deployment of large clusters or supercomputing environments. Background working with fast-moving AI labs or frontier-model infrastructure deployments and customer-facing roles (field engineering, or pre/post-sales architecture). Systems engineering, coding, and debugging skills including experience with C/C++, Linux kernel, and drivers. Hands-on experience with NVIDIA GPU systems and SDKs (e.g., CUDA), NVIDIA networking technologies (NICs, RoCE, InfiniBand), and/or ARM-based CPU solutions. Familiarity with virtualization and cloud-native networking concepts. We make extensive use of conferencing tools, but occasional travel (up to 20%) is required for on-site customer visits and industry