Software Engineer, Infrastructure - Analytics Platform

at OpenAI
USD 230,000-385,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

CI/CD @ 3 ClickHouse @ 3 Distributed Systems @ 3 Kubernetes @ 3 Networking @ 3 Observability @ 3 Profiling @ 3 Prometheus @ 3 Python Rust @ 6 Terraform @ 3

Details

The Scaling team designs, builds, and operates critical infrastructure that enables research at OpenAI. Its systems range from low-level infrastructure components to research-facing custom applications and must scale with increasingly complex workloads while remaining reliable and easy to use.

The role involves designing, developing, and operating production-critical software infrastructure across its full production lifecycle. It focuses on backend and systems engineering, including low-level performance, scalability, distributed systems, and hands-on development and operation of critical services at scale.

This is not a general Python backend role. The position requires strong systems experience in Rust or C++, particularly in performance-sensitive infrastructure. The role is based in San Francisco, California, with a hybrid work model requiring three days in the office per week.

Responsibilities

  • Own critical infrastructure across design, implementation, rollout, operation, and iteration.
  • Build and operate performant backend systems in Rust or C++ that support core research workflows.
  • Design and improve distributed data and serving systems, including partitioning, replication, consistency, retries, backpressure, and failure isolation.
  • Debug production bottlenecks involving latency, throughput, contention, hot spots, and overload behavior.
  • Operate business-critical services through on-call rotations, incident response, postmortems, observability, alerting, safe rollouts, rollback plans, and zero-downtime migrations.
  • Improve the reliability of services running on Kubernetes, including resource tuning, failure handling, and production readiness.
  • Partner with engineers and researchers to deliver fast, reliable, and useful systems.
  • Raise technical standards through sound judgment, ownership, and follow-through.

Requirements

  • A track record of owning operationally critical systems end to end and delivering outcomes in ambiguous environments.
  • Strong hands-on experience building performance-sensitive backend systems in Rust or C++.
  • Experience with concurrency, asynchronous execution, memory behavior, serialization, I/O, networking, profiling, and failure analysis.
  • Experience designing, building, or operating distributed systems or distributed databases at meaningful scale.
  • Preferably, experience with ClickHouse-like systems or infrastructure for analytics, telemetry, logging, search, ingestion, storage, or query execution.
  • Hands-on experience with production infrastructure and delivery tooling, including Kubernetes, Terraform or equivalent infrastructure-as-code tooling, Prometheus or equivalent observability and metrics tooling, and CI/CD pipelines or equivalent release automation.
  • Hands-on experience operating production-critical systems across the full lifecycle, including incidents, observability, alerting, safe rollouts, rollback plans, and recurrence prevention.
  • Strong judgment in balancing engineering quality, speed, risk, and business impact.
  • A habit of shipping practical first versions and improving them through production feedback.

Benefits

  • Base salary of $230,000–$385,000 per year.
  • Equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health and dependent care expenses, as well as commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits may be provided, including charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer committed to providing reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs