Technical Lead Manager - Training Runtime, Data(set) Movement

at OpenAI
USD 295,000-445,000 per year
SENIOR
✅ Hybrid
✅ Relocation

Tech Stack

API @ 7 Data Pipelines @ 4 Debugging @ 7 Distributed Systems @ 4 Machine Learning @ 4 Python @ 6 Reinforcement Learning @ 4 Rust @ 4

Details

Training Runtime builds the distributed systems that power OpenAI's largest model training runs. The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time and moves model state safely and efficiently across large clusters.

The work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. The goal is to enable researchers to move quickly while keeping training runs fast, reproducible, debuggable, and resilient at scale.

Responsibilities

  • Design and build a unified dataset read platform for multiple current and future training frameworks.
  • Define dataset APIs, storage-format expectations, registration and versioning, and migration paths that make data access reproducible and maintainable.
  • Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts.
  • Build terminal and web-based visualizers for inspecting text, multimodal, and reinforcement learning data late in the pipeline.
  • Write and review production code in core data loading, service, caching, and reliability paths.
  • Partner with teams working on training frameworks, reinforcement learning, multimodal models, storage, runtime, and cluster infrastructure.
  • Initially serve as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners.
  • Over time, expand ownership to broader data movement systems, including checkpoint loads and saves and snapshot transfers.

Requirements

  • Experience building or owning dataset, data loading, storage, or distributed training infrastructure at large scale, such as torch.utils.data.
  • Strong understanding of API design, debugging ergonomics, performance, and bit-level correctness.
  • Understanding of failure modes in large distributed training jobs and how data systems can create or prevent them.
  • Experience with stateful iterators, checkpoint and restart semantics, caching, remote services, or high-throughput storage reads.
  • Comfort working across Python and lower-level systems code. Rust or C++ experience is useful but not required.
  • Experience with multimodal, video, reinforcement learning, or pretraining data pipelines.
  • Ability to lead through code and technical judgment before a team exists, and later manage engineers without losing a hands-on focus.
  • Strong focus on developer experience, including eliminating manual preprocessing scripts and niche cluster-specific bugs.

Benefits

  • Base salary of $295,000–$445,000 per year, plus equity.
  • Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs