Cpu/Storage/Pop-Wan Program Manager

at OpenAI
USD 226,000-285,000 per year
MIDDLE SENIOR
✅ Hybrid
✅ Relocation

Tech Stack

AI Azure Communication @ 6 GPU Leadership @ 6 Networking @ 6

Details

OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical.

The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly.

Responsibilities

  • Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint
  • Drive readiness to convert contracted compute capacity into schedulable production clusters
  • Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
  • Build integrated schedules spanning procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff
  • Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
  • Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation
  • Manage deployment of storage systems supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans
  • Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers
  • Lead physical deployment execution including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria
  • Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale
  • Identify risks early across supply chain, site readiness, technical constraints, and vendor execution, then drive mitigation plans
  • Communicate milestones, escalations, and capacity forecasts to senior leadership

Requirements

  • 8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
  • Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
  • Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling
  • Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors
  • Strong understanding of deployment lifecycles from planning and procurement through production handoff
  • Ability to reason across physical infrastructure execution and logical systems architecture dependencies
  • Proven ability to build integrated schedules and drive accountability across multiple stakeholders
  • Strong executive communication skills with experience managing critical escalations and leadership updates
  • Comfortable operating in fast-moving environments with aggressive timelines and evolving priorities
  • Highly analytical with strong problem-solving and execution instincts

Benefits

  • Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
  • Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
  • 401(k) retirement plan with employer match
  • Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
  • Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
  • 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
  • Mental health and wellness support
  • Employer-paid basic life and disability coverage
  • Annual learning and development stipend to fuel your professional growth
  • Daily meals in our offices, and meal delivery credits as eligible
  • Relocation support for eligible employees
  • Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided

More details about our benefits are available to candidates during the hiring process.

More jobs at OpenAI

Similar jobs