Hardware Technical Program Manager, Infrastructure Partner Operations

at OpenAI
USD 226,000-285,000 per year
SENIOR
✅ On-site
✅ Relocation

Tech Stack

AI @ 3 AWS @ 4 Azure @ 4 Communication @ 6 Distributed Systems @ 7 GCP @ 4 GPU @ 3 HPC @ 3 Leadership @ 6 Networking Oracle @ 4 Reporting @ 4

Details

The Industrial Compute team builds the physical infrastructure powering OpenAI's largest-scale AI systems. The team designs, deploys, and operates next-generation compute infrastructure across a global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners.

The role leads operational delivery across third-party infrastructure partners, including major cloud service providers and strategic compute vendors. It serves as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous improvement. The role coordinates with Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership.

Responsibilities

  • Own operational engagement with third-party infrastructure providers, ensuring execution against operational commitments, service-level agreements (SLAs), and performance expectations.
  • Develop operational governance frameworks, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
  • Define, track, and improve metrics covering infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.
  • Build dashboards and reporting mechanisms to provide visibility into partner performance, risks, trends, and areas requiring executive attention.
  • Coordinate internal teams and external providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
  • Lead escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance.
  • Establish operating rhythms with partners, including weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives.
  • Partner with Capacity Planning, Hardware Operations, Networking, Deployment, Reliability Engineering, and Supply Chain teams to align providers with operational priorities.
  • Identify systemic operational risks and drive corrective actions that improve long-term operational effectiveness.

Requirements

  • 7+ years of experience in Technical Program Management, Infrastructure Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments.
  • Experience managing operational relationships with external infrastructure providers, cloud service providers, hardware vendors, or strategic technology partners.
  • Strong understanding of hyperscale cloud infrastructure, data center operations, infrastructure delivery, or large-scale distributed systems.
  • Experience developing operational KPIs, SLAs, service health metrics, dashboards, and executive reporting.
  • Demonstrated success leading cross-functional operational programs involving internal stakeholders and external partners.
  • Strong program management skills and the ability to drive accountability across organizations without direct authority.
  • Excellent written and verbal communication skills, including presenting operational performance to senior technical and executive leadership.
  • Bachelor's degree in Engineering, Computer Science, Information Systems, Operations, or equivalent practical experience.

Preferred Skills

  • Experience managing cloud infrastructure operations with Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), or other hyperscale cloud providers.
  • Experience with operational governance, service delivery, customer success engineering, technical account management, or infrastructure operations for enterprise cloud customers.
  • Strong understanding of SLAs, operational KPIs, incident management, escalation processes, root cause analysis, and continuous service improvement methodologies.
  • Experience building executive dashboards, operational scorecards, business review frameworks, and data-driven performance reporting.
  • Familiarity with GPU infrastructure, AI infrastructure, high-performance computing (HPC), or hyperscale data center environments.
  • Experience managing complex cross-company technical relationships while balancing customer priorities, engineering constraints, and operational execution.
  • Ability to influence senior stakeholders across internal teams and external partner organizations without direct authority.
  • Experience driving continuous operational improvements through metrics, process optimization, and structured governance.

Benefits

  • Base salary range of $226,000–$285,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax FSA, dependent care FSA, and commuter accounts.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, and paid office closures.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs