Software Engineer, Data Infrastructure

at OpenAI
USD 185,000-385,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 3 Airflow @ 3 Distributed Systems @ 3 Flink @ 3 Kafka @ 3 Machine Learning Security Spark @ 3 Terraform @ 3 Trino @ 3

Details

Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. The team operates large-scale Spark compute fleets, builds data lakes and metadata systems on Iceberg and Delta, runs high-throughput streaming platforms on Kafka and Flink, provides orchestration with Airflow, and supports ML feature engineering tooling such as Chronon.

The role focuses on building and operating high-performance, scalable data infrastructure. This includes scaling and hardening big data compute and storage platforms, supporting high-throughput streaming systems, operating low-latency data ingestion, enabling secure and governed data access for ML and analytics, and designing for reliability and performance at extreme scale. The role includes full lifecycle ownership spanning architecture, implementation, production operations, and on-call participation.

This position is exclusively based at OpenAI's San Francisco headquarters and follows a hybrid work model with three days in the office per week.

Responsibilities

  • Design, build, and maintain data infrastructure systems, including distributed compute, data orchestration, distributed storage, streaming infrastructure, and machine learning infrastructure.
  • Ensure scalability, reliability, efficiency, and security across data platform systems.
  • Scale the data platform by orders of magnitude while maintaining reliability and efficiency.
  • Empower engineers and teammates with excellent data tooling and systems.
  • Collaborate with product, research, and analytics teams to build the technical foundations that unlock new features and experiences.
  • Own the reliability of systems, including participation in an on-call rotation for critical incidents.

Requirements

  • 4+ years of experience in data infrastructure engineering, or 4+ years of infrastructure engineering experience with a strong interest in data.
  • Experience supporting Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms.
  • Experience with infrastructure tooling such as Terraform.
  • Ability to debug large-scale distributed systems.
  • Interest in solving data infrastructure problems in the AI space.
  • Experience building and operating scalable, reliable, and secure systems.
  • Comfort with ambiguity and rapid change.
  • Intrinsic desire to learn and fill in missing skills, along with the ability to share learnings clearly and concisely.

Benefits

  • Base salary range of $185,000–$385,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health and dependent care expenses, as well as commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and paid sick or safe time as required by applicable law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer committed to providing reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs