Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 3
Communication @ 3
Distributed Systems @ 6
GPU @ 3
HPC @ 3
Kubernetes @ 3
Linux @ 3
Networking
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. The team builds Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters.
Responsibilities
- Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
- Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers.
- Build provisioning and configuration services coordinating network boot, hardware management interfaces, firmware, operating-system images, drivers, and host configuration.
- Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
- Design reliable reconciliation and recovery for concurrent changes, interrupted operations, and partial failures, using staged rollouts to limit disruption across nodes, racks, and clusters.
- Improve control-plane throughput, API latency, and the time infrastructure takes to reach its desired state while respecting site-system and provider-API limits.
- Build software integrations for new sites and GPU hardware generations, partnering with hardware, networking, data-center, and other infrastructure teams.
Requirements
- Strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.
- Experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.
- Understanding of how a bare-metal node moves from power-on to a configured, workload-ready system, with depth in one or more of PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.
- Ability to design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.
- Ability to diagnose reliability and performance problems across service, operating-system, and machine boundaries and turn production evidence into lasting software improvements.
- Effective collaboration across engineering specialties and clear communication of system behavior and technical tradeoffs.
Bonus Qualifications
- Experience building infrastructure control planes coordinating operations across multiple sites or regions.
- Experience with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters.
- Experience integrating multiple hardware platforms or infrastructure providers into a common service or resource model.
Benefits
- Base salary range of $255,000–$490,000 per year, plus equity and performance-related bonuses for eligible employees.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts, 401(k) with employer match, paid parental and caregiver leave, paid time off, paid holidays, mental health and wellness support, and employer-paid basic life and disability coverage.
- Annual learning and development stipend, daily office meals, meal delivery credits where eligible, and relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.
More jobs at OpenAI
Field Enablement Lead, Technical Success
OpenAI · San Francisco, United States
USD 272,000-302,000 per year
Staff PLM & Engineering Applications Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 257,000-285,000 per year
Business Systems Lead, Procure-to-Pay
OpenAI · San Francisco, United States
USD 272,000-302,000 per year
Software Engineer, Healthcare
OpenAI · San Francisco, United States
USD 347,000-385,000 per year
Software Engineer, Manufacturing Infrastructure
OpenAI · San Francisco, United States
USD 310,000-460,000 per year
Similar jobs
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Technical Product Manager - GenAI Platforms - AI Infrastructure - CTO Office
Bloomberg · New York City, United States
USD 140,000-295,000 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Nvidia · United States
USD 152,000-287,500 per year
Applied Systems Engineering Rotation Engineer - New College Graduate 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Nvidia · Santa Clara, United States
USD 160,000-322,000 per year
Principal Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year