Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Azure
CUDA @ 2
Debugging @ 3
Distributed Systems @ 3
GPU
HPC
InfiniBand @ 2
MPI @ 2
Machine Learning @ 3
NCCL @ 2
NVLink @ 2
PyTorch @ 2
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Team
The Inference team brings OpenAI's most capable research and technology to the world through its products. The team enables consumers, enterprise customers, and developers to use advanced AI models, focusing on performant and efficient model inference and accelerating research through model inference.
About the Role
The role involves optimizing large-scale AI models for high-volume, low-latency, and high-availability production and research environments.
Responsibilities
- Collaborate with machine learning researchers, engineers, and product managers to bring new technologies into production.
- Work with researchers to enable advanced research through engineering.
- Introduce techniques, tools, and architectures that improve the performance, latency, throughput, and efficiency of the model inference stack.
- Build tools to identify bottlenecks and sources of instability, then design and implement solutions for high-priority issues.
- Optimize code and Azure virtual machine fleets to fully utilize GPU compute and GPU RAM.
Requirements
- Understanding of modern machine learning architectures and how to optimize their performance, particularly for inference.
- At least 5 years of professional software engineering experience.
- Familiarity with or ability to quickly learn PyTorch, NVIDIA GPUs, and optimization software stacks such as NCCL and CUDA.
- Familiarity with high-performance computing technologies such as InfiniBand, MPI, and NVLink.
- Experience architecting, building, observing, and debugging production distributed systems, preferably performance-critical systems.
- Experience rebuilding or substantially refactoring production systems to support rapidly increasing scale.
- Ability to own problems end-to-end, work independently, and learn as needed.
- A collaborative, humble attitude and willingness to help colleagues and contribute to team success.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics.
Background checks may be administered in accordance with applicable law. OpenAI is committed to providing reasonable accommodations to applicants with disabilities.
Benefits
- Base pay range of $295,000–$555,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time as required by law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.