Staff Technical Program Manager, AI Infrastructure
General Motors
- Location
- Sunnyvale, California, United States of America
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 267 approvals (FY2023)
- Posted
- Aug 21, 2026
Skills
About this role
Job Description
Staff Technical Program Manager, AI Infrastructure At General Motors, our product teams are redefining mobility. Through a human-centered design process, we create vehicles and experiences that are designed not just to be seen, but to be felt. We’re turning today’s impossible into tomorrow’s standard – from breakthrough hardware and battery systems to intuitive design, intelligent software, and next-generation safety and entertainment features. Every day, our products move millions of people as we aim to make driving safer, smarter, and more connected, shaping the future of transportation on a global scale.
The Role
We are seeking a Staff Technical Program Manager (TPM) to lead AV ML Infrastructure programs for our autonomous driving platform. In this role, you will own strategy and execution for large-scale ML infrastructure – including training pipelines, model lifecycle management, compute orchestration, and platform reliability – that power next-generation autonomy models. You will operate at the intersection of ML engineering, platform infrastructure, and operations, ensuring our systems are scalable, efficient, and production-ready to support end-to-end model development at scale. What You’ll Do Program Leadership: Own end-to-end delivery of ML infrastructure programs, driving measurable improvements in training throughput, platform reliability, and developer productivity. Establish clear goals, milestones, and success metrics across teams. Cross-Functional Alignment: Partner with ML engineers, platform teams, validation, and product to prioritize initiatives, drive tradeoff decisions, and accelerate the AI development lifecycle. Technical Roadmapping: Translate complex MLOps challenges – distributed training orchestration, compute scheduling, pipeline scaling – into clear, actionable plans with defined ownership and outcomes. Scalability & Reliability: Drive infrastructure evolution to support growing model complexity, dataset scale, and compute demand, with a strong focus on resiliency, observability, and performance. Risk & Dependency Management: Identify risks early, manage cross-team dependencies, and implement mitigation strategies to ensure stable, predictable delivery. Operational Excellence: Establish best practices for monitoring, incident response, and capacity planning to ensure high system uptime and efficient resource utilization. Metrics & Visibility: Define and track KPIs (e.g., system reliability, utilization, training cycle time), delivering clear, executive-ready insights on program health and progress. Your Skills & Abilities (Required Qualifications) 10+ years of technical program management experience leading large, complex, cross-functional initiatives 5+ years working in ML infrastructure, MLOps, AI platform engineering, or distributed compute environments BS or MS in Engineering, Computer Science, or a related technical field Experience delivering large-scale ML infrastructure programs, including compute orchestration, pipeline reliability, and resource management Proven ability to lead programs spanning infrastructure, software, and data systems in ambiguous, fast-evolving environments Strong analytical skills with the ability to interpret system metrics and drive performance improvements Excellent communication and stakeholder management skills, with the ability to influence across technical and non-technical audiences Deep familiarity with Agile delivery, JIRA (or similar tools), and technical program reporting frameworks What will give you a competitive edge (Preferred Qualifications) Experience scaling large-scale ML infrastructure, including GPU compute, cluster orchestration (e.g., Kubernetes, Slurm), or cloud platforms (AWS, GCP, Azure) Familiarity with ML workflow orchestration and MLOps tooling (e.g., Kubeflow, Airflow) Background in SRE, platform engineering, or DevOps practices applied to distributed ML systems Experience with observability