Technical Lead Manager - Training Runtime, Data(set) Movement
at OpenAI
USD 295,000-445,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 7
Data Pipelines @ 4
Debugging @ 7
Distributed Systems @ 4
Machine Learning @ 4
Python @ 6
Reinforcement Learning @ 4
Rust @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Training Runtime builds the distributed systems that power OpenAI's largest model training runs. The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time and moves model state safely and efficiently across large clusters.
The work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. The goal is to enable researchers to move quickly while keeping training runs fast, reproducible, debuggable, and resilient at scale.
Responsibilities
- Design and build a unified dataset read platform for multiple current and future training frameworks.
- Define dataset APIs, storage-format expectations, registration and versioning, and migration paths that make data access reproducible and maintainable.
- Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts.
- Build terminal and web-based visualizers for inspecting text, multimodal, and reinforcement learning data late in the pipeline.
- Write and review production code in core data loading, service, caching, and reliability paths.
- Partner with teams working on training frameworks, reinforcement learning, multimodal models, storage, runtime, and cluster infrastructure.
- Initially serve as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners.
- Over time, expand ownership to broader data movement systems, including checkpoint loads and saves and snapshot transfers.
Requirements
- Experience building or owning dataset, data loading, storage, or distributed training infrastructure at large scale, such as
torch.utils.data. - Strong understanding of API design, debugging ergonomics, performance, and bit-level correctness.
- Understanding of failure modes in large distributed training jobs and how data systems can create or prevent them.
- Experience with stateful iterators, checkpoint and restart semantics, caching, remote services, or high-throughput storage reads.
- Comfort working across Python and lower-level systems code. Rust or C++ experience is useful but not required.
- Experience with multimodal, video, reinforcement learning, or pretraining data pipelines.
- Ability to lead through code and technical judgment before a team exists, and later manage engineers without losing a hands-on focus.
- Strong focus on developer experience, including eliminating manual preprocessing scripts and niche cluster-specific bugs.
Benefits
- Base salary of $295,000–$445,000 per year, plus equity.
- Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Forward Deployed Engineer (FDE), Legal-SF
OpenAI · New York City, United States, San Francisco, United States
USD 162,000-280,000 per year
Software Engineer, Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 210,000-405,000 per year
Manager, Applied AI Engineering (Codex)
OpenAI · New York City, United States, San Francisco, United States
USD 251,000-335,000 per year
Product Design Manager, Growth - Consumer
OpenAI · San Francisco, United States
USD 347,000-405,000 per year
Marketing Scientist
OpenAI · New York City, United States, San Francisco, United States
USD 234,000-260,000 per year
Similar jobs
Machine Learning Engineer, API Multicloud
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Member of Technical Staff (AI Inference Engineer)
Perplexity AI · New York City, United States, Palo Alto, United States, San Francisco, United States
USD 220,000-485,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
AI Systems Engineer, Codex Agents
OpenAI · San Francisco, United States
USD 230,000-385,000 per year