Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Azure
Communication @ 7
Compliance
GPU @ 4
HPC @ 4
Leadership @ 7
Networking @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. The CPU / Storage / PoP / WAN team is responsible for server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion across first-party facilities, colocation environments, and strategic cloud partners.
The highly technical Program Manager will lead execution across CPU, Storage, PoP, and WAN infrastructure programs that unlock next-generation compute capacity. The role owns complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion, coordinating hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies.
This role is based in San Francisco, California, with travel as needed.
Responsibilities
- Lead end-to-end execution of CPU/GPU cluster activation programs across OpenAI’s global infrastructure footprint.
- Drive readiness to convert contracted compute capacity into schedulable production clusters.
- Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives.
- Build integrated schedules covering procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff.
- Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones.
- Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation.
- Manage storage system deployments supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans.
- Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers.
- Lead physical deployment execution, including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria.
- Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale.
- Identify risks across supply chain, site readiness, technical constraints, and vendor execution, and drive mitigation plans.
- Communicate milestones, escalations, and capacity forecasts to senior leadership.
Requirements
- 8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations.
- Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems.
- Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling.
- Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors.
- Strong understanding of deployment lifecycles from planning and procurement through production handoff.
- Ability to reason across physical infrastructure execution and logical systems architecture dependencies.
- Proven ability to build integrated schedules and drive accountability across multiple stakeholders.
- Strong executive communication skills, including experience managing critical escalations and leadership updates.
- Comfort operating in fast-moving environments with aggressive timelines and evolving priorities.
- Strong analytical, problem-solving, and execution skills.
Preferred Skills
- Experience at a hyperscaler, cloud provider, AI infrastructure company, or global network operator.
- Experience deploying GPU clusters, HPC systems, or large training environments.
- Familiarity with distributed storage systems and high-performance data infrastructure.
- Experience with PoP deployments, WAN backbone expansion, or global network buildouts.
- Experience working across first-party, colocation, and cloud environments.
- Experience building repeatable infrastructure deployment systems in high-growth environments.
Benefits
- Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer committed to reasonable accommodations and compliance with applicable employment laws.