Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 3
Data Pipelines @ 6
Debugging @ 3
Distributed Systems @ 6
GPU @ 3
LLM
Machine Learning @ 3
Mathematics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Workload team designs and operates OpenAI's LLM training and inference infrastructure, providing scalable systems for researchers working across large GPU and accelerator fleets.
The role focuses on designing and implementing dataset infrastructure for OpenAI's next-generation training stack. This includes standardized dataset interfaces, scalable pipelines, performance testing, and collaboration with multimodal researchers and infrastructure teams.
Responsibilities
- Design and maintain standardized dataset APIs, including APIs for multimodal data that cannot fit in memory.
- Build proactive testing and scale-validation pipelines for dataset loading at GPU scale.
- Integrate datasets into training and inference pipelines to support smooth adoption and user experience.
- Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt.
- Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged after standardization.
- Debug and resolve performance bottlenecks in distributed dataset loading, including straggler systems that slow global training.
- Provide visualization and inspection tools to identify dataset errors, bugs, and bottlenecks.
Requirements
- Strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
- Experience building APIs, modular code, and scalable abstractions.
- Understanding that abstractions serve users and that user experience is an important part of abstraction design.
- Comfort debugging bottlenecks across large fleets of machines.
- Interest in building reliable, scalable infrastructure and owning a foundational part of the machine learning stack.
- Collaborative and humble working style.
Bonus Qualifications
- Background knowledge in data mathematics, probability, or distributed data theory.
- Experience with GPU-scale distributed systems or dataset scaling for real-time data.
Benefits
- Base salary range of $250,000–$380,000 per year, plus equity.
- Medical, dental, and vision insurance for employees and families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.
More jobs at OpenAI
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Android Systems Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 216,000-342,000 per year
Senior Staff Software Engineer, Identity
OpenAI · Mountain View, United States, San Francisco, United States
USD 345,000-405,000 per year
Analytics Engineer, GTM
OpenAI · San Francisco, United States, New York City, United States
USD 220,000-335,000 per year
Product Designer, Payments
OpenAI · San Francisco, United States
USD 245,000-310,000 per year
Similar jobs
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Member of Technical Staff (AI Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior AI Systems and Algorithms Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year