Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Azure
Communication @ 6
GPU
Leadership @ 6
Networking @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical.
The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly.
Responsibilities
- Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint
- Drive readiness to convert contracted compute capacity into schedulable production clusters
- Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
- Build integrated schedules spanning procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff
- Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
- Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation
- Manage deployment of storage systems supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans
- Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers
- Lead physical deployment execution including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria
- Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale
- Identify risks early across supply chain, site readiness, technical constraints, and vendor execution, then drive mitigation plans
- Communicate milestones, escalations, and capacity forecasts to senior leadership
Requirements
- 8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
- Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
- Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling
- Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors
- Strong understanding of deployment lifecycles from planning and procurement through production handoff
- Ability to reason across physical infrastructure execution and logical systems architecture dependencies
- Proven ability to build integrated schedules and drive accountability across multiple stakeholders
- Strong executive communication skills with experience managing critical escalations and leadership updates
- Comfortable operating in fast-moving environments with aggressive timelines and evolving priorities
- Highly analytical with strong problem-solving and execution instincts
Benefits
- Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
- 401(k) retirement plan with employer match
- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
- 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
- Mental health and wellness support
- Employer-paid basic life and disability coverage
- Annual learning and development stipend to fuel your professional growth
- Daily meals in our offices, and meal delivery credits as eligible
- Relocation support for eligible employees
- Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided
More details about our benefits are available to candidates during the hiring process.
More jobs at OpenAI
Hardware Technical Program Manager, Infrastructure Partner Operations
OpenAI · San Francisco, United States
USD 226,000-285,000 per year
Software Engineer, Conversion Measurement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Customer Learning Program Lead
OpenAI · San Francisco, United States
USD 266,000-300,000 per year
AI Deployment Engineer, Agent Enablement
OpenAI · San Francisco, United States
USD 197,000-280,000 per year
Software Engineer, Ads Integrity
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Similar jobs
Principal Security Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 347,000-490,000 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Principal Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 347,000-490,000 per year
Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Security Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Technical Program Manager, Compute
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 290,000-365,000 per year
Director, Compute Infrastructure Procurement Operations
Anthropic · San Francisco, United States, New York City, United States
USD 300,000-385,000 per year