Software Engineer, Compute Infrastructure

at OpenAI
USD 230,000-405,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 3 Communication @ 3 Distributed Systems @ 3 GPU @ 3 HPC Kubernetes @ 3 NCCL @ 3 Networking @ 3 Observability @ 3 Performance Optimization @ 3 Profiling @ 3 Security

Details

Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. The team designs, provisions, schedules, operates, and optimizes systems connecting accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability.

The work spans capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, systems optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and infrastructure developer experience.

Responsibilities

  • Build and deeply optimize reliable system software for large-scale compute systems running demanding AI workloads.
  • Design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health.
  • Profile, benchmark, and optimize training workloads across compute, memory, storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks.
  • Create hardware-aware automation for provisioning, firmware and driver upgrades, incident response, and day-to-day operations.
  • Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools for researchers, product engineers, and operators.
  • Turn operational lessons into better systems, stronger abstractions, and clearer ownership boundaries.
  • Collaborate across research, engineering, security, networking, hardware, and data center teams.

Potential areas of work include compute foundations, fleet and orchestration, core network engineering, hardware health and observability, storage, and agent infrastructure.

Requirements

  • Strong software engineering skills and experience building, operating, or improving production infrastructure systems.
  • Experience in one or more of distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high-performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware-aware performance optimization, benchmarking, developer experience, or infrastructure tooling.
  • Ability to debug complex system behavior across software, hardware, networking, and workload layers and turn findings into robust improvements.
  • Comfort with ambiguity, strong ownership, and a bias toward practical, durable solutions.
  • Interest in infrastructure that directly enables frontier AI research and product impact.
  • Clear communication and the ability to work effectively across teams with different constraints and goals.

Benefits

  • Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
  • Pre-tax FSA, dependent care FSA, and commuter accounts.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, and paid office closures.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Equity, performance-related bonuses for eligible employees, and additional taxable fringe benefits.

OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.

More jobs at OpenAI

Similar jobs