Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API
ChatGPT
Communication @ 3
Debugging @ 3
Docker
Go
Kafka
Kubernetes
Observability
PostgreSQL
Python
Rust
Terraform
TypeScript
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Engineering Acceleration team builds and operates the foundational systems engineers use to build, test, and ship ChatGPT, the API, and OpenAI's infrastructure.
This role focuses on evolving OpenAI's build and continuous integration systems for a fast-growing engineering organization. The work spans developer productivity, build systems, distributed infrastructure, and software quality, including Bazel-based builds, Buildkite pipelines, test selection, remote caching and execution, CI observability, and tooling for diagnosing failures.
Responsibilities
- Own and evolve Bazel-based build and test workflows across a large, polyglot monorepo.
- Design and maintain Starlark rules, macros, toolchains, and integrations that make builds reproducible, hermetic, and easy for product teams to adopt.
- Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, test sharding, retry behavior, and flake isolation.
- Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling.
- Improve local development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly.
- Operate and optimize build infrastructure across Docker and OCI images, Kubernetes-based runners, cloud resources, and remote cache and execution systems.
- Instrument build and CI systems with metrics, logs, traces, dashboards, and analytics to measure speed, reliability, cost, and developer impact.
- Partner with product, infrastructure, and research engineering teams to understand pain points, onboard projects, debug build issues, and remove systemic bottlenecks.
- Use modern AI tools for CI failure analysis, flaky test debugging, pull request triage, automatic remediation, and developer-facing explanations.
- Own the reliability of the systems built, including participation in an on-call rotation for critical developer infrastructure.
Technologies
- Bazel and Starlark
- Buildkite
- Docker and OCI images
- Kubernetes
- Python, Go, TypeScript, Rust, C++, and other languages
- Terraform
- Remote caching and remote execution
- Artifact storage and build telemetry systems
- Postgres, Kafka, and internal engineering platform services
Requirements
- 5+ years of software engineering experience, including significant experience building infrastructure or tooling for developers.
- Hands-on experience with Bazel, Buck, Pants, Gradle, or similar build systems.
- Understanding of hermetic builds, dependency graphs, caching, sandboxing, and remote execution.
- Experience building or operating CI systems at scale, particularly where build time, queue time, test flakiness, and developer trust affect engineering velocity.
- Ability to write production software for internal platforms, including code, design, debugging, operations, and long-term ownership.
- Ability to debug distributed build and CI failures across source control, dependency management, containers, runners, remote caches, test frameworks, and service infrastructure.
- Strong focus on developer experience and reducing operational toil.
- Pragmatic approach to platform adoption and building faster, clearer, and more reliable paved paths.
- Clear communication across teams and ability to turn ambiguous productivity problems into concrete technical plans.
- Interest in applying AI to developer infrastructure without weakening quality, reliability, or safety.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics. Background checks and reasonable accommodation processes are administered in accordance with applicable law.
Benefits
- Salary range of $185,000–$490,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid holidays, office closures, and sick or safe time as applicable.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and eligible meal delivery credits.
- Relocation support for eligible employees.
- Potential additional benefits including charitable donation matching and wellness stipends.