Technical Program Manager, Model Deployment & Capacity
OpenAI
- Location
- San Francisco
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 2h ago
About this role
About the Team
The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches.
About the Role
We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner with product, research, inference, fleet, and capacity teams to develop scenarios, surface tradeoffs, and drive timely allocation decisions. Lead model deployment readiness and rollout planning, including serving-capacity allocation, launch sequencing, validation, and operational handoffs. Establish clear readiness gates, risk reviews, rollback criteria, and escalation paths for model deployments. Drive launch coordination through deployment and post-launch learning, turning recurring gaps and manual work into scalable tooling and operating practices. Define and operationalize metrics for forecast accuracy, capacity utilization and headroom, deployment velocity, reliability, latency, quality, and user impact. Create concise, decision-ready communications that make dependencies, risks, capacity constraints, and launch choices clear to technical and product leaders. You might thrive in this role if you: Have led complex technical programs in infrastructure, distributed systems, capacity planning, model serving, or large-scale deployment environments. Can reason credibly about demand, supply, headroom, reliability, latency, and quality tradeoffs, and translate them into executable plans. Have built operating mechanisms or tooling that replaced fragmented, manual workflows with scalable systems and clear ownership. Are effective in high-ambiguity, constrained environments where priorities change and decisions require explicit tradeoffs. Build alignment across research, engineering, product, finance or capacity planning, and operations without relying on direct authority. Use metrics to guide decisions, identify bottlenecks, and demonstrate measurable improvements in throughput, predictability, or reliability. Communicate with precision and can move comfortably between technical detail, operational execution, and executive-level decisions. Thrive in ambiguous, scaling environments and can bring order to complex cross-functional work without losing pace. Care