Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Debugging @ 2
GPU @ 6
Machine Learning
Networking
Performance Analysis @ 3
Profiling @ 2
System Architecture @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI's Infrastructure organization builds and evaluates systems that power advanced AI workloads. The Scaling team works across hardware, modeling, and architecture to understand workload behavior across evolving hardware platforms and connect theoretical capability with observed system performance.
The role focuses on evaluating new hardware platforms by porting benchmarks and real-world workloads, analyzing performance, identifying system bottlenecks, and adapting workloads to better use hardware capabilities. The position is based in San Francisco, California, follows a hybrid model with three days in the office per week, and offers relocation assistance.
Responsibilities
- Port and enable benchmarks and real-world workloads on new hardware platforms.
- Evaluate system performance across compute, memory, storage, and networking subsystems.
- Identify and analyze performance bottlenecks and inefficiencies.
- Adapt and optimize workloads to better utilize hardware capabilities.
- Develop and run performance experiments and profiling workflows.
- Compare expected and observed performance and provide feedback to hardware architecture, performance modeling, system, and software engineering teams.
- Debug issues across the stack, including software, runtime, and hardware interactions.
- Provide actionable insights to guide platform readiness and deployment decisions.
Requirements
- Experience with performance analysis, benchmarking, or workload optimization.
- Strong understanding of system architecture, including CPU/GPU, memory, and I/O subsystems.
- Experience porting or adapting workloads across different hardware platforms.
- Familiarity with profiling tools and performance debugging techniques.
- Ability to identify root causes of performance issues across hardware and software boundaries.
- Experience working in large-scale or distributed system environments.
Preferred Skills
- Experience with AI/ML workloads, including training or inference systems.
- Familiarity with GPU or accelerator-based systems.
- Experience with low-level performance tools, including profilers, tracing, and microbenchmarks.
- Background in systems software, compilers, or runtime optimization.
- Experience collaborating with hardware and architecture teams on performance validation.
Benefits
- Base salary of $293,000–$385,000 per year.
- Equity and performance-related bonus opportunities for eligible employees.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and paid sick or safe time as required by law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Software Engineer, AI Accelerator Runtime
OpenAI · San Francisco, United States
USD 266,000-445,000 per year
People Research Scientist
OpenAI · Mountain View, United States, San Francisco, United States
USD 198,000-220,000 per year
Systems Software Engineer, Silicon Bringup
OpenAI · San Francisco, United States
USD 266,000-445,000 per year
Software Engineer, Model Runtime
OpenAI · San Francisco, United States
USD 266,000-445,000 per year
Strategic Delivery Lead, Intelligence Community
OpenAI · Washington, United States
USD 266,000-370,000 per year
Similar jobs
Principal System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Performance Modeling Engineer
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Performance Modeling Engineer ~2
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year