Training Performance Engineer

at OpenAI
USD 250,000-445,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

CUDA @ 1 Communication @ 2 Distributed Systems @ 3 GPU @ 3 HPC @ 3 JAX @ 3 MPI @ 2 NCCL @ 2 Observability Performance Analysis Profiling @ 3 PyTorch @ 3 Python @ 1 Rust @ 1 TensorFlow @ 3

Details

Training Runtime designs the core distributed machine-learning training runtime that powers early research experiments through frontier-scale model runs. The team is building a unified, modular runtime focused on high-performance data movement, fault-tolerant training frameworks, deterministic orchestration, observability, and distributed process management.

The role drives efficiency improvements across the distributed training stack by analyzing large-scale training runs, identifying utilization gaps, and designing optimizations that improve throughput and uptime. This includes GPU kernel performance analysis, collective communication throughput, I/O bottleneck investigation, and model sharding for massive-scale training. The role is based in San Francisco, California, with a hybrid work model requiring three days in the office per week.

Responsibilities

  • Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage.
  • Optimize GPU utilization and throughput for large-scale distributed model training.
  • Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
  • Implement model graph transformations to improve end-to-end throughput.
  • Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
  • Partner with researchers to ensure new model architectures scale efficiently during pre-training.
  • Contribute to infrastructure decisions that improve the reliability and efficiency of large training jobs.

Requirements

  • Strong programming skills in Python and C++; Rust or CUDA experience is a plus.
  • Experience running distributed training jobs on multi-GPU systems or HPC clusters.
  • Ability to debug complex distributed systems and measure efficiency rigorously.
  • Exposure to frameworks such as PyTorch, JAX, or TensorFlow.
  • Understanding of how large-scale training loops are built.
  • Ability to collaborate across teams and translate profiling data into practical engineering improvements.

Nice to Have

  • Familiarity with NCCL, MPI, or UCX communication libraries.
  • Experience with large-scale data loading and checkpointing systems.
  • Prior work on training runtimes, distributed scheduling, or machine-learning compiler optimization.

Benefits

  • Base salary range of $250,000–$445,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs