Technical Program Manager, Storage & Data Infrastructure

at OpenAI
USD 257,000-445,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

API @ 3 Azure @ 6 Communication @ 6

Details

About the Team

OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems support data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history, and snapshots, the underlying architecture increasingly shapes product capabilities, latency, reliability, cost, and efficiency.

About the Role

We are looking for a technically deep Technical Program Manager to independently define and lead multiple programs across data platforms, online databases, and storage infrastructure. You will connect model, product, and data-consumer requirements to architecture and work with engineering teams to take new capabilities through production adoption and repeatable expansion.

The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The role focuses on making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment, and validation.

This role is based in San Francisco, California, and follows a hybrid work model with three days in the office per week. Relocation assistance is available to new employees.

Responsibilities

  • Translate model, product, and data-platform needs into precise access-pattern, consistency, durability, freshness, availability, and scalability requirements.
  • Connect memory, history, retrieval, and resumable work to product capability and end-to-end latency.
  • Partner with engineering to transform data and storage architecture into repeatable scale units, including standardized provisioning, placement, routing, data movement, and readiness checks.
  • Lead cross-stack programs spanning ingestion and processing, databases and indexes, and file/object storage.
  • Define data ownership, schema compatibility, change-data-capture, replay, and consumer-readiness contracts.
  • Evaluate physical versus logical footprint, index and replication amplification, redundant copies, tiering, caching, and network movement in relation to cost and workload value.
  • Drive resilience and recovery programs covering database backup, failover, point-in-time recovery, execution/workspace save-and-restore, recovery time, safe resumption, and isolation from live traffic.
  • Coordinate lifecycle correctness across files, objects, databases, and data platforms, including metadata, retention, deletion, and snapshots.
  • Incorporate privacy, access control, auditability, and residency requirements into designs and consumer contracts.
  • Lead adoption and major migrations through compatibility checks, representative workload testing, staged cutovers, rollback, and operational handoff.
  • Improve APIs, guardrails, and self-service capabilities so new capacity and capabilities can be consumed predictably.
  • Measure architecture changes through product and platform outcomes such as task completion, continuity, data freshness, query and retrieval latency, snapshot latency, throughput, reliability, and cost efficiency.

Requirements

  • Experience independently owning complex production programs in data platforms, databases, or storage infrastructure, with the ability to explain architectural decisions, personal contributions, and resulting impact.
  • Deep working knowledge of hyperscaler and cloud storage technologies, such as Amazon S3 or Azure Blob Storage, including performance, placement, resiliency, and cost constraints.
  • Understanding of the full data and storage stack, including product access patterns and APIs; ingestion and processing; databases, storage engines, and indexes; caching, replication, and data movement; and durable-storage, CPU, and network layers.
  • Ability to translate model, product, and data-consumer needs into precise platform requirements and define evidence that capabilities, recovery paths, and scaling changes are ready.
  • Experience leading delivery across product, model, data, database, storage, and infrastructure teams, particularly when no single team owns the end-to-end result.
  • Ability to use workload evidence to make tradeoffs across capability, correctness, reliability, latency, and cost efficiency.
  • Strong communication skills and the ability to maintain ownership through production adoption, not only launch milestones.

Benefits

  • Base salary of $257,000–$445,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax FSA, dependent care, and commuter accounts.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs