Software Engineer, Infrastructure - Core Experimentation

at OpenAI
USD 293,000-325,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI ChatGPT @ 3 Codex @ 3 Debugging Distributed Systems @ 3 Experimentation Observability @ 6 Performance Monitoring

Details

The Statsig team within OpenAI owns the experimentation, rollout, dynamic configuration, and analytics infrastructure that supports OpenAI product launches. The platform helps teams ship safely, evaluate product and model changes in production, and make high-confidence decisions from real-world usage.

The team builds infrastructure for ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and shared platform systems. Its systems support configuration evaluation, traffic management, experiment-data ingestion, analytics, and safe rollout or rollback of production changes.

As an Infrastructure Engineer on the Statsig team, you will build and scale the foundational systems behind OpenAI's experimentation and rollout platform. You will work on the distributed control plane, SDK and server evaluation paths, ingestion pipelines, analytics foundations, and operational tooling that make launches safe and measurable at OpenAI scale.

Responsibilities

  • Design and operate low-latency configuration delivery systems powering feature flags, dynamic configurations, and progressive rollouts across OpenAI product suites.
  • Scale SDK, server-side evaluation, and control-plane systems so high-volume services can depend on Statsig without adding user-visible latency or operational fragility.
  • Build high-throughput data ingestion and analytics infrastructure for experimentation, product analytics, feature performance monitoring, and model or product measurement workflows.
  • Improve the performance, efficiency, reliability, and observability of core Statsig infrastructure as OpenAI products scale globally.
  • Optimize query performance, data freshness, and data availability for teams making launch decisions from experimentation and analytics workflows.
  • Strengthen operational excellence through improved SLOs, alerting, debugging tools, incident response, capacity planning, and failure-mode design.
  • Partner with teams across ChatGPT, Codex, model measurement, consumer ads, business subscriptions, developer products, and infrastructure to turn recurring launch and measurement needs into durable platform capabilities.
  • Lead large technical initiatives and shape the architecture of experimentation and rollout infrastructure used across the company.

Requirements

  • Experience building large-scale distributed systems with strict performance, availability, and correctness requirements.
  • Experience with low-latency systems work, including real-time configuration delivery, SDK or runtime performance, caching, concurrency, and high-throughput service design.
  • Experience building or operating large-scale data platforms, event pipelines, analytics systems, or query infrastructure where freshness and correctness matter.
  • Strong focus on reliability, observability, incident response, capacity planning, and operating critical production systems.
  • Ability to take ownership of complex technical problems end to end and build infrastructure that enables other teams to move faster with confidence.
  • Interest in preserving Statsig's platform strengths while integrating deeply into OpenAI's product and platform stack.

Location

This role is based in Bellevue, Washington. The team works in person three days per week and offers relocation assistance to new employees.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs