Principal Software Engineer, Simulation

at OpenAI
USD 347,000-490,000 per year
SENIOR
✅ On-site
✅ Relocation

Tech Stack

API @ 6 Codex LLM Python @ 6 Rust @ 1

Details

About the Team

OpenAI's research training infrastructure powers how its frontier models are trained and evaluated. The Simulation team sits at the intersection of the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that the model's training environment is as realistic as possible.

The team owns the integration layer that connects production harness capabilities into the training stack. Researchers depend on it to run experiments and evaluations reliably and to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness.

About the Role

The Principal Software Engineer will lead the architecture and evolution of the Simulation Platform. The role owns a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments.

This role is suited to a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work involves building robust infrastructure that supports and accelerates research without compromising engineering quality.

Responsibilities

  • Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next.
  • Build a platform for LLMs to train and be evaluated in simulated environments that closely mimic their deployment settings across the agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and other dimensions.
  • Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance.
  • Build reliable execution systems that support demanding training workloads at scale.
  • Partner closely with research, agent, infrastructure, and platform teams to support new training use cases and harness capabilities.
  • Design clean, stable interfaces and workflows for highly technical internal users who move quickly and expect strong ergonomics.
  • Prevent one-off workarounds from becoming long-term technical debt by establishing durable abstractions and clear ownership.
  • Raise the bar for correctness, reliability, operational rigor, and engineering judgment across a critical research-facing system.

Requirements

  • Significant experience building and scaling backend or infrastructure systems in fast-moving environments.
  • Deep strength in API design, systems design, and engineering fundamentals.
  • Strong attention to detail and a deep commitment to correctness, reliability, and operational quality.
  • Ability to work directly with demanding technical users while maintaining strong engineering discipline.
  • A track record of leading cross-functional technical efforts and creating clarity across organizational boundaries.
  • Strong product sense and user empathy for internal platforms and developer tooling.
  • Motivation to enable researchers and accelerate their work rather than conducting research directly.
  • Proficiency in Python and experience with backend platform engineering.
  • Rust experience is a plus.

Location

This role is ideally based in San Francisco due to the close collaboration required with researchers and applied engineering partners.

Benefits

  • Base pay of $347,000–$490,000 per year, in addition to equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health and dependent care expenses, as well as commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and paid sick or safe time as required by applicable law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer and is committed to providing reasonable accommodations to applicants with disabilities. Background checks will be administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs