Training: ML Framework Engineer

at OpenAI
USD 205,000-445,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 6 Distributed Systems GPU Machine Learning @ 6 Observability Performance Optimization Python @ 5

Details

Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. The team is building a unified, modular runtime that accelerates researchers and enables frontier-scale training.

The work focuses on high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, highly available, fault-tolerant training frameworks, including training loops, state management, resilient checkpointing, deterministic orchestration, and observability; and distributed process management for long-lived, job-specific, and user-provided processes.

The team integrates large-scale capabilities into a composable, developer-facing runtime, partnering closely with model-stack, research, and platform teams. Success is measured by improving both training throughput and researcher throughput.

As a Training: ML Framework Engineer, you will improve training throughput for the internal training framework while enabling researchers to experiment with new ideas. This includes designing, implementing, and optimizing state-of-the-art AI models, writing reliable machine-learning code, and developing a deep understanding of supercomputer performance. The role focuses on performance optimization, distributed systems, and highly reliable code. The training framework is used for large runs involving massive numbers of GPUs.

This role is based in San Francisco, California, and follows a hybrid work model with three days in the office per week.

Responsibilities

  • Apply the latest techniques in the internal training framework to achieve high hardware efficiency for training runs.
  • Profile and optimize the training framework.
  • Work with researchers to enable development of next-generation models.

Requirements

  • Experience running small-scale machine-learning experiments.
  • Strong interest in understanding how systems work and improving their speed while minimizing complexity and maintenance burden.
  • Strong software engineering skills.
  • Proficiency in Python.

Benefits

  • Base salary of $205,000–$445,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and paid sick or safe time as applicable.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.
  • OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.

More jobs at OpenAI

Similar jobs