Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 3
CUDA
Communication @ 1
GPU @ 3
Machine Learning
Networking
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Workload Networking team is responsible for the collective communication stack used in the company's largest training jobs. Using C++ and CUDA, the team develops novel collective communication techniques that enable efficient training of flagship models on custom-built supercomputers.
As a Software Engineer, Networking, you will design and implement custom networking collectives tightly integrated into the training stack. The role focuses on low-level, performance-critical software, with collective communication experience considered a bonus.
Responsibilities
- Collaborate closely with machine learning researchers to design and implement efficient collective operations in C++ and CUDA.
- Ensure that the largest training jobs take full advantage of the different network transports used in the company's supercomputers.
- Develop simulations to inform future supercomputer network designs.
Requirements
- Experience writing distributed algorithms using RDMA.
- Comfort writing low-level, performance-sensitive CPU and/or GPU code.
- Familiarity with network simulation techniques.
- Experience with collective communication is a bonus.
Work Arrangement
- Based in San Francisco, California.
- Hybrid work model with three days in the office per week.
- Relocation assistance is offered to eligible new employees.
Benefits
- Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses, as well as commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may include charitable donation matching and wellness stipends.
The base pay range is $380,000–$555,000 per year. Total compensation also includes equity and performance-related bonuses for eligible employees. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.