Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
GPU
Machine Learning @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to advance OpenAI's mission of building AGI that benefits all of humanity.
As a Technical Lead, you will work with machine learning researchers and design talent to advance model capabilities. This role is based in San Francisco, California, follows a hybrid model with three days per week in the office, and offers relocation assistance to eligible employees.
Responsibilities
- Evaluate and select silicon platforms, including GPUs, NPUs, and specialized accelerators, for on-device and edge deployment of OpenAI models.
- Work with research teams to co-design model architectures that meet real-world deployment constraints, including latency, memory, power, and bandwidth.
- Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities.
- Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads.
- Build and lead a team responsible for implementing the low-level inference stack, including kernel development and runtime systems.
- Turn emerging research capabilities into capabilities that can support further development.
Requirements
- Experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators.
- Understanding of transformer model performance characteristics, including attention, KV-cache behavior, and memory bandwidth requirements.
- Experience designing or optimizing high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware machine learning pipelines.
- Experience building or leading teams working on low-level, performance-critical software such as CUDA kernels, compilers, or machine learning runtimes.
- Experience working extensively with models involving language and perception.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company pushes the boundaries of AI system capabilities and seeks to safely deploy them through its products.
OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics. Background checks are administered in accordance with applicable law. Reasonable accommodations are available to applicants with disabilities.