Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
CI/CD @ 6
ChatGPT
Communication @ 6
GitHub @ 3
Kubernetes @ 3
Networking @ 2
Observability
Python @ 3
Security @ 2
TypeScript @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The ChatGPT Velocity team owns the health of the end-to-end developer and deployment loop for ChatGPT-facing services. Its mission is to help ChatGPT engineers ship confidently through a fast, lightweight development cycle and a trusted path to production.
The team works across local environments, Bazel builds, CI, tests, merges, release candidates, deployments, synthetics, alerts, and production health to reduce development friction and improve reliability.
Responsibilities
- Reduce local development startup and iteration latency for ChatGPT engineers.
- Improve Bazel and build performance for local machines, devboxes, and CI, including cache effectiveness, reproducibility, disk usage, remote cache behavior, and observability into invalidations.
- Improve CI reliability and merge throughput by reducing unrelated failures, tail latency, flaky or mis-owned tests, quarantine friction, and repeated manual reruns.
- Build agents that help pull requests make forward progress through rebasing, conflict resolution, CI triage, safe reruns, owner routing, and identification of external blockers.
- Make deployment pipelines more self-healing, particularly when transient alerts, known-bad clusters, recovered synthetics, stale release candidates, or noisy gates do not require human intervention.
- Improve deployment observability so engineers can quickly determine where a change is, what is blocking it, whether the signal is trustworthy, and who owns the next action.
- Establish lightweight metrics for impact, including commit-to-production latency, deployment-window throughput, CI tail latency, local bootup time, failure rates, manual interventions, false-positive gates, and developer-reported pain.
Requirements
- Strong experience with large-scale developer infrastructure, build systems, CI/CD, deployment systems, or production reliability.
- Ability to debug across local machines, devboxes, remote caches, monorepos, service dependencies, authentication systems, and distributed deployment pipelines.
- Comfortable working with Bazel, Buildkite, Kubernetes, Temporal, Python, TypeScript services, GitHub workflows, and large monorepo tooling, or able to ramp quickly on equivalent systems.
- Product-oriented systems thinking, with a focus on whether developers can understand and trust workflows in addition to whether subsystems are technically healthy.
- Ability to replace repeated manual intervention with policy-backed automation while preserving production safety.
- Strong cross-team communication skills and the ability to turn ambiguous reports into clear owners, hypotheses, metrics, and shipped fixes.
- Preference for practical measurement while recognizing developer frustration as an early signal of systems problems.
- Interest in high-leverage work where small improvements can save significant engineering time.
Nice to Have
- Experience improving developer experience in a very large monorepo.
- Experience with remote execution, remote caching, hermetic builds, or cache reproducibility.
- Experience designing CI quarantine, test ownership, merge queue, or auto-revert systems.
- Experience with progressive delivery, canary, preview, or stable deployment pipelines; synthetics; alert quality; and deployment-manager workflows.
- Experience building agent-assisted developer workflows, such as automated pull request monitoring, CI triage, or conflict resolution.
- Familiarity with security and access-control constraints involving secrets, internal authentication, VPN or exit-node networking, and local development.
Benefits
- Equity, performance-related bonuses for eligible employees, and benefits in addition to base pay.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.
More jobs at OpenAI
Strategic Delivery Lead, Cyber
OpenAI · Washington, United States
USD 266,000-370,000 per year
Technical Program Manager, Global Programs — Applied AI Engineering
OpenAI · San Francisco, United States
USD 251,000-280,000 per year
Software Engineer, DevOps
OpenAI · San Francisco, United States
USD 177,000-327,000 per year
Program Manager, Government Trusted Access
OpenAI · Washington, United States
USD 162,000-240,000 per year
Senior Web Experience Strategist
OpenAI · San Francisco, United States
USD 287,000-318,000 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff Site Reliability Engineer - AI Platform Runtime
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Security Engineer, Application Security
Sentry · San Francisco, United States
USD 155,000-400,000 per year
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Nvidia · Santa Clara, United States
USD 160,000-322,000 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year