Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 1
Debugging @ 3
Distributed Systems @ 3
GPU
HPC
Kubernetes @ 3
LLM
Machine Learning
NCCL @ 3
NVLink
Networking @ 2
Performance Optimization
Profiling @ 6
PyTorch @ 3
Python @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Scaling team is responsible for the architectural and engineering backbone of OpenAI's infrastructure, designing and delivering systems that support the deployment and operation of advanced AI models. Its work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization.
The role focuses on enabling production workloads and end-to-end testing on new platforms. Responsibilities include creating test harnesses and platform stress benchmarks, porting inference and training workloads to new and early-access systems or hardware, analyzing performance bottlenecks, and characterizing end-to-end system behavior across compute, communications, storage, control planes, and failure modes.
Responsibilities
- Port and validate key inference and training workloads on new platforms and SKUs, driving correctness, performance, and stability to an internal readiness standard.
- Build benchmarks and stress tests that capture real end-to-end workload behavior across CPUs, GPUs, memory subsystems, frontends, scale-up and scale-out networking, WAN traffic, NVLink and RDMA collectives, storage, thermals, and other relevant system components.
- Analyze distributed training and inference performance, including:
- Collective performance and tuning across NCCL, RCCL, and internal libraries.
- Compute and communication overlap.
- Kernel-level bottlenecks.
- Memory bandwidth and scheduling effects.
- Create repeatable test harnesses for CI and lab environments that produce actionable outputs, including pass/fail results, performance scores, and regression detection.
- Partner with systems and fleet bring-up engineers to ensure platforms are stable, performant, operationally usable, and scalable, including containerization, Kubernetes integration, telemetry hooks, and failure-triage processes.
- Work cross-functionally with vendors and internal stakeholders by producing clear bug reports, minimal reproductions, and prioritized issue lists.
Requirements
- Bachelor's degree in computer science, electrical engineering, or equivalent practical experience.
- Five or more years of experience in one or more of the following areas: ML systems, performance engineering, distributed systems, or high-performance computing.
- Hands-on experience with PyTorch and modern large language model training and inference stacks.
- Understanding of large-scale distributed training concepts, including data, model, and pipeline parallelism and collective communications.
- Experience with RDMA and debugging or optimizing communications libraries such as NCCL or RCCL, including their interaction with hardware and networks.
- Proficiency in Python and comfort reading or writing performance-critical code. C++, CUDA, or HIP experience is a plus.
- Strong profiling and debugging skills, including tools such as Nsight, rocprof, perf, and flamegraphs, with the ability to reason from traces and performance counters.
Preferred Skills
- Experience building workload-shaped benchmarks and stress or fault tests that correlate with production behavior rather than only synthetic loops or microbenchmarks.
- Familiarity with RDMA networking and transport tuning, including how network topology and congestion affect collectives.
- Experience running and validating workloads in Kubernetes and bridging research code into robust, repeatable infrastructure.
- Hands-on laboratory experience with early hardware, including new network interface cards, GPUs or accelerators, and early racks.
Benefits
- Equity, performance-related bonuses for eligible employees, and benefits in addition to the listed base salary.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax Flexible Spending Accounts and commuter benefits.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may be provided, including charitable donation matching and wellness stipends.
- OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.
More jobs at OpenAI
Support Program Manager, Partnerships
OpenAI · San Francisco, United States
USD 216,000-240,000 per year
Researcher, Recursive Self-Improvement Safety
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Manager, Cyber - AI Deployment Engineering
OpenAI · San Francisco, United States
USD 302,000-335,000 per year
Support Program Manager, Live Support
OpenAI · San Francisco, United States
USD 216,000-240,000 per year
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year