Software Engineer, Workload Enablement

at OpenAI
USD 293,000-385,000 per year
MIDDLE SENIOR
✅ Hybrid
✅ Relocation

Tech Stack

AI CUDA @ 1 Debugging @ 3 Distributed Systems @ 3 GPU HPC Kubernetes @ 3 LLM Machine Learning NCCL @ 3 NVLink Networking @ 2 Performance Optimization Profiling @ 6 PyTorch @ 3 Python @ 5

Details

The Scaling team is responsible for the architectural and engineering backbone of OpenAI's infrastructure, designing and delivering systems that support the deployment and operation of advanced AI models. Its work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization.

The role focuses on enabling production workloads and end-to-end testing on new platforms. Responsibilities include creating test harnesses and platform stress benchmarks, porting inference and training workloads to new and early-access systems or hardware, analyzing performance bottlenecks, and characterizing end-to-end system behavior across compute, communications, storage, control planes, and failure modes.

Responsibilities

  • Port and validate key inference and training workloads on new platforms and SKUs, driving correctness, performance, and stability to an internal readiness standard.
  • Build benchmarks and stress tests that capture real end-to-end workload behavior across CPUs, GPUs, memory subsystems, frontends, scale-up and scale-out networking, WAN traffic, NVLink and RDMA collectives, storage, thermals, and other relevant system components.
  • Analyze distributed training and inference performance, including:
    • Collective performance and tuning across NCCL, RCCL, and internal libraries.
    • Compute and communication overlap.
    • Kernel-level bottlenecks.
    • Memory bandwidth and scheduling effects.
  • Create repeatable test harnesses for CI and lab environments that produce actionable outputs, including pass/fail results, performance scores, and regression detection.
  • Partner with systems and fleet bring-up engineers to ensure platforms are stable, performant, operationally usable, and scalable, including containerization, Kubernetes integration, telemetry hooks, and failure-triage processes.
  • Work cross-functionally with vendors and internal stakeholders by producing clear bug reports, minimal reproductions, and prioritized issue lists.

Requirements

  • Bachelor's degree in computer science, electrical engineering, or equivalent practical experience.
  • Five or more years of experience in one or more of the following areas: ML systems, performance engineering, distributed systems, or high-performance computing.
  • Hands-on experience with PyTorch and modern large language model training and inference stacks.
  • Understanding of large-scale distributed training concepts, including data, model, and pipeline parallelism and collective communications.
  • Experience with RDMA and debugging or optimizing communications libraries such as NCCL or RCCL, including their interaction with hardware and networks.
  • Proficiency in Python and comfort reading or writing performance-critical code. C++, CUDA, or HIP experience is a plus.
  • Strong profiling and debugging skills, including tools such as Nsight, rocprof, perf, and flamegraphs, with the ability to reason from traces and performance counters.

Preferred Skills

  • Experience building workload-shaped benchmarks and stress or fault tests that correlate with production behavior rather than only synthetic loops or microbenchmarks.
  • Familiarity with RDMA networking and transport tuning, including how network topology and congestion affect collectives.
  • Experience running and validating workloads in Kubernetes and bridging research code into robust, repeatable infrastructure.
  • Hands-on laboratory experience with early hardware, including new network interface cards, GPUs or accelerators, and early racks.

Benefits

  • Equity, performance-related bonuses for eligible employees, and benefits in addition to the listed base salary.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax Flexible Spending Accounts and commuter benefits.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time as required by applicable law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits may be provided, including charitable donation matching and wellness stipends.
  • OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.

More jobs at OpenAI

Similar jobs