Software Engineer, Productivity - Inference Runtime

at OpenAI
USD 230,000-385,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

AI API CI/CD @ 3 ChatGPT Codex Debugging @ 3 Distributed Systems @ 3 GPU Python @ 3

Details

We’re hiring a Developer Productivity Engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads.

This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack.

Responsibilities

  • Improve systems that ensure inference engine releases are correct, performant, and regression-free by evolving tooling and infrastructure for deploy gate validation.
  • Bring rigor to release, validation, branching, and deployment processes across the inference stack.
  • Improve canary, asynchronous, and large-scale validation workflows for inference systems.
  • Harden CI, testing, and validation infrastructure so failures are actionable and trustworthy.
  • Reduce noisy or flaky failures caused by infrastructure instability, GPU scheduling, or test environment issues.
  • Build automation for failure triage, ownership detection, debugging, and escalation.
  • Partner with inference teams, research developer productivity, engine acceleration, and infrastructure teams to improve release quality and rollout safety.
  • Reduce developer friction in testing, debugging, and release workflows so engineers can move faster with confidence.

A major focus will be improving tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across metrics such as time-to-first-token (TTFT) and time-between-tokens (TBT).

Requirements

  • Strong experience with CI/CD systems, testing infrastructure, release tooling, developer productivity, or large-scale build and validation systems.
  • Comfort working in Python-heavy environments and debugging complex distributed systems.
  • Strong developer empathy and an interest in improving workflows, reducing friction, and making engineers more effective.
  • High ownership, including the ability to proactively identify problems, drive improvements, and follow issues through resolution.
  • Interest in building automation that reduces manual triage, improves signal quality, and scales operational effectiveness.
  • Comfort operating in ambiguous areas without a fully predefined roadmap.
  • Pragmatic, collaborative, and motivated by helping teams move faster with more confidence.
  • Technical curiosity and willingness to learn about large-scale inference systems; prior inference experience is not required.

Python experience is highly relevant because much of the current deploy gate and validation infrastructure is Python-based. C++ experience is helpful, particularly for working near inference engine code, CI build issues, or performance-sensitive systems, but it is not required.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics.

Background checks will be administered in accordance with applicable law. OpenAI is committed to providing reasonable accommodations to applicants with disabilities.

Benefits

  • Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, and paid office closures.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily meals in offices and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits may also be provided.

The base pay offered may vary depending on market location, job-related knowledge, skills, and experience. Total compensation also includes generous equity and performance-related bonuses for eligible employees.

More jobs at OpenAI

Similar jobs