Tech Stack
AI API ChatGPT GPU Networking SecurityDetails
The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for next-generation models. The team operates a large, modern GPU fleet and provides a unified platform for production Applied AI and Research training workloads.
This role owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. The role spans hardware, networking, power, and cooling, driving execution, risk management, and alignment from working teams through leadership to deliver production-ready capacity at scale.
The role is based in San Francisco, California, with a hybrid work model requiring three days in the office per week.
Responsibilities
- Lead end-to-end delivery of new compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference.
- Drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling, owning plans, dependencies, and critical paths.
- Interface with chip providers to reduce risks related to onboarding new hardware platforms by working across kernel, communications, hardware, and scheduling engineering teams.
- Build and operationalize program mechanisms, including roadmaps, milestones, risk registers, and runbooks, to make delivery predictable at massive scale.
- Partner with engineering to improve cluster turn-up reliability, repeatability, and automation, reducing time to serve new capacity.
- Support network operations and the end-to-end physical and logical bring-up of OpenAI network Points of Presence, including on-site deployment and rack cabling.
- Coordinate cross-functional readiness across security, finance, operations, and product and research stakeholders to ship production-ready compute.
- Manage integrations and handoffs across teams and partners, ensuring consistent execution, clear communication, and fast issue resolution.
- Identify bottlenecks and systemic gaps, then drive durable fixes across tooling, processes, and partner interfaces.
- Provide executive visibility into progress, trade-offs, and risks across a large portfolio of concurrent programs.
Requirements
- A degree in a hard science or a demonstrated track record of engineering expertise.
- Five or more years of program management experience for major projects, including capital projects or hyperscaler infrastructure deployments.
- Ability to serve as the primary person responsible for driving and delivering complex projects.
- Experience managing cross-functional and cross-company teams, including driving information and decision hygiene.
- An extensive track record of successfully delivering high-profile technical projects against tight deadlines.
- Technical aptitude and experience partnering effectively with engineering or fundamental research teams.
- Experience interfacing with and leading external vendors, including engineering firms, equipment suppliers, and/or construction firms.
- Expertise designing and implementing simple, scalable processes that solve complex problems.
- Experience managing complicated dependencies such as logistics and/or supply chains.
- Resourcefulness and the ability to thrive in ambiguous, fast-paced environments.
- Interest in and thoughtful consideration of the impacts of AGI.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company develops AI systems and seeks to deploy them safely, with safety and human needs at the core.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.