Software Engineer, Compute Foundations

at OpenAI
USD 255,000-490,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

API @ 3 Communication @ 3 Distributed Systems @ 6 GPU @ 3 HPC @ 3 Kubernetes @ 3 Linux @ 3 Networking

Details

Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. The team builds Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters.

Responsibilities

  • Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
  • Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers.
  • Build provisioning and configuration services coordinating network boot, hardware management interfaces, firmware, operating-system images, drivers, and host configuration.
  • Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
  • Design reliable reconciliation and recovery for concurrent changes, interrupted operations, and partial failures, using staged rollouts to limit disruption across nodes, racks, and clusters.
  • Improve control-plane throughput, API latency, and the time infrastructure takes to reach its desired state while respecting site-system and provider-API limits.
  • Build software integrations for new sites and GPU hardware generations, partnering with hardware, networking, data-center, and other infrastructure teams.

Requirements

  • Strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.
  • Experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.
  • Understanding of how a bare-metal node moves from power-on to a configured, workload-ready system, with depth in one or more of PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.
  • Ability to design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.
  • Ability to diagnose reliability and performance problems across service, operating-system, and machine boundaries and turn production evidence into lasting software improvements.
  • Effective collaboration across engineering specialties and clear communication of system behavior and technical tradeoffs.

Bonus Qualifications

  • Experience building infrastructure control planes coordinating operations across multiple sites or regions.
  • Experience with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters.
  • Experience integrating multiple hardware platforms or infrastructure providers into a common service or resource model.

Benefits

  • Base salary range of $255,000–$490,000 per year, plus equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts, 401(k) with employer match, paid parental and caregiver leave, paid time off, paid holidays, mental health and wellness support, and employer-paid basic life and disability coverage.
  • Annual learning and development stipend, daily office meals, meal delivery credits where eligible, and relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs