Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 5
Algorithms
Debugging @ 3
Distributed Systems
LLM
Machine Learning
PyTorch
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for advanced AI workloads. The Accelerators team evaluates and brings up new compute platforms that support large-scale AI training and inference.
The role involves prototyping system software on new accelerators, enabling performance optimizations across AI workloads, and working across kernels, sharding strategies, distributed systems scaling, and performance modeling. You will help adapt OpenAI’s software stack to non-traditional hardware and improve the efficiency of core AI workloads at scale. This is not a compiler-focused role; it bridges ML algorithms with system performance.
Responsibilities
- Prototype and enable OpenAI’s AI software stack on new and exploratory accelerator platforms.
- Optimize the performance of large-scale models, including LLMs, recommender systems, and distributed AI workloads, across diverse hardware environments.
- Develop kernels, sharding mechanisms, and system scaling strategies for emerging accelerators.
- Collaborate on optimizations at the model code level, such as PyTorch, and below to improve performance on non-traditional hardware.
- Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization.
- Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures.
- Contribute to runtime improvements, compute/communication overlapping, and scaling efforts for frontier AI workloads.
Requirements
- 3+ years of experience working on AI infrastructure, including kernels, systems, or hardware-software co-design.
- Hands-on experience with data-center-scale AI accelerator platforms, such as TPUs, custom silicon, or exploratory architectures.
- Strong understanding of kernels, sharding, runtime systems, or distributed scaling techniques.
- Familiarity with optimizing LLMs, CNNs, or recommender models for hardware efficiency.
- Experience with performance modeling, system debugging, and software stack adaptation for novel architectures.
- Ability to operate across multiple levels of the stack, rapidly prototype solutions, and navigate ambiguity during early hardware bring-up phases.
- Exposure to mobile accelerators is welcome, though experience enabling data-center-scale AI hardware is preferred.
- Interest in shaping the future of AI compute through exploration of alternatives to mainstream accelerators.
- Candidates may need to meet certain legal status requirements under U.S. export control laws and regulations.
Benefits
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses, parking, and transit.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time as required by law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.