Software Engineer - Data Aquisition (Systems)

at OpenAI
USD 255,000-405,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

CI/CD @ 3 Debugging Distributed Systems @ 3 Git @ 3 Kubernetes @ 3 Linux @ 3 Networking @ 3 Observability @ 3

Details

About the Team

This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly.

The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel.

About the Role

As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability—while also writing software to make these systems better over time.

This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work.

Responsibilities

  • Build and operate reliable infrastructure for research workloads and research-facing services.
  • Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services.
  • Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
  • Debug issues across networking, compute, storage, orchestration, and service reliability layers.
  • Build software and automation that reduce manual operational work and improve system reliability.
  • Partner closely with researchers, infrastructure engineers, and service owners to understand system needs and translate them into durable solutions.
  • Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
  • Take ownership of critical systems and drive work independently from problem definition through execution.

Requirements

You might thrive in this role if you:

  • Have strong systems fundamentals and understand how infrastructure scales in practice.
  • Are comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.
  • Can write software to automate, debug, and improve infrastructure systems.
  • Have a strong execution mindset and can independently drive ambiguous infrastructure work.
  • Enjoy supporting a wide surface area of systems, from research tooling to platform services.
  • Are pragmatic about when to build custom systems versus using existing, well-supported tools.
  • Care about building reliable systems that make researchers faster and reduce operational friction.

Nice to have:

  • Experience with PXE boot, cluster provisioning, bare-metal infrastructure, or large-scale fleet management.
  • Experience operating Kubernetes or similar orchestration systems at scale.
  • Experience with infrastructure-as-code, CI/CD, observability, or deployment automation.
  • Experience supporting search infrastructure, data platforms, ingest systems, or large-scale research workflows.
  • Experience with Git-based workflows and internal developer tooling.
  • Prior experience in environments where reliability, scale, and speed all matter.

Benefits

  • Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
  • Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
  • 401(k) retirement plan with employer match
  • Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
  • Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
  • 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
  • Mental health and wellness support
  • Employer-paid basic life and disability coverage
  • Annual learning and development stipend to fuel your professional growth
  • Daily meals in our offices, and meal delivery credits as eligible
  • Relocation support for eligible employees
  • Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.

More details about our benefits are available to candidates during the hiring process.

More jobs at OpenAI

Similar jobs