Software Engineer, Manufacturing Infrastructure
OpenAI
- Location
- San Francisco
- Employment
- Full Time
- Work model
- Remote
- Level
- Mid
- Sponsorship
- Sponsors visa
- Posted
- 1h ago
Skills
About this role
About the Team
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.
About the Role
We're seeking a Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate the software used to manufacture, qualify, and deliver our hardware from the factory. You will work across the stack to create the infrastructure that runs internally and externally to coordinate all aspects of the production process. You will create the critical tools and procedures to execute, capture, process, and present the data resulting from the end to end assembly and validation of our hardware across multiple vendors and sites. This role is hands-on and high-ownership. You will work closely across teams both internal and external to define the standards that will be used across our products to ensure the velocity and quality of our 1P hardware. You will own the implementation, deployment, and output of these systems as well their continued maintenance and SLAs. Location: San Francisco, CA (Hybrid: 3 days/week onsite). Relocation assistance available. In this role, you will: Design, develop, and maintain the software infrastructure for manufacturing process execution and data export. Own integration across internal customers and vendor systems and processes. Build and maintain the CI, release, and delivery pipeline of tooling to external partners. Build and maintain internal systems to ingest, process, deliver, and visualize critical data for internal teams and systems. Build system health monitoring, telemetry, remote diagnostics, and recovery paths that make manufacturing failures traceable and actionable. Develop automation frameworks and monitoring for board bring-up, rack bring-up, qualification, manufacturing readiness, deployment readiness, and long-term reliability. Package engineering releases into manufacturing-ready software recipes: images, versions, logs, limits, remediation mapping, provisioning hooks, secure artifact handling, and traceable data export. Debug complex production issues spanning hardware, firmware, kernel, services, networking, power, thermals, boot, provisioning, manufacturing, test, and cloud services. Partner with hardware, firmware, security, networking, infrastructure, manufacturing, operations, and external engineering teams to define software contracts, unblock bring-up, and drive issues to closure. Produce durable architecture notes, runbooks, validation records, and decision documents that help OpenAI and partner teams reproduce, operate, and improve the platform. You might thrive in this role if you have: 7+ years of hands-on experience, or exceptional accomplishments demonstrating equivalent expertise, in system software, platform software, hardware diagnostics, and cloud infrastructure. Strong programming skills in C, C++, Rust or similar systems languages, with experience building reliable software for real hardware. Familiarity with scripting in languages such as shell and python as used in automation systems. Familiarity with Linux-based hardware platforms, OpenBMC, Redfish, firmware update systems, kernel drivers, and fleet management software. Demonstrated ability to debug live hardware using logs, packet captures, firmware traces, bus captures, lab hosts, BMC journals, Linux tooling, and carefully controlled experiments. Experience with hardware bring-up, manufacturing or qualification testing, system diagnostics, release validation, or deployment of high-performance compute, accelerator, server, networking, storage, or embedded