Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
CUDA @ 3
Debugging @ 4
Distributed Systems @ 7
GPU
Go @ 7
LLVM @ 3
Linux @ 7
Machine Learning @ 4
Networking @ 7
Observability
Profiling @ 4
Python @ 7
Rust @ 7
SGLang @ 4
Security
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI is building infrastructure that enables advanced AI models to be deployed reliably, efficiently, and at global scale. The GPT Infrastructure team develops software that turns inference and optimization research into production products across distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships.
The role focuses on building an automated inference optimization platform. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes. You will design both the OpenAI-hosted control plane and partner-side software that evaluates candidates on real accelerator hardware.
Responsibilities
- Design, build, and operate durable APIs and control-plane services for multi-hour or multi-day optimization campaigns, including scheduling, retries, budgets, checkpoints, artifact lineage, and observability.
- Build secure partner-side runner and grader software that can compile, execute, verify, and benchmark candidate artifacts on third-party accelerator hardware.
- Integrate hardware profiles, ISA and toolchain context, compilers, runtimes, and inference-serving engines into repeatable optimization workflows.
- Turn research prototypes into reliable product surfaces with clear contracts, debuggable failure modes, reproducible outputs, and excellent developer ergonomics.
- Develop correctness and performance evaluation systems spanning latency, throughput, memory use, utilization, and cost efficiency.
- Build artifact, provenance, and qualification workflows that make optimized kernels, binaries, configurations, and reports safe to review and deploy.
- Collaborate with Research, Inference Engineering, Infrastructure, Security, Product, and Strategic Partnerships to deliver production-ready solutions.
- Drive technical architecture and execution across ambiguous, cross-functional initiatives connecting OpenAI systems with partner environments.
Requirements
- 8+ years of professional software engineering experience building large-scale distributed systems, infrastructure platforms, or cloud services, or equivalent depth of experience.
- Strong programming skills in one or more of C++, Python, Go, or Rust.
- Experience designing and operating highly available backend systems, APIs, job orchestration systems, or durable workflows for production workloads.
- Strong understanding of distributed systems, Linux, networking, storage, containers, and modern cloud architectures.
- Experience debugging complex systems and using measurement, profiling, and benchmarks to guide engineering decisions.
- Proven ability to lead complex technical initiatives as a senior individual contributor and work effectively across organizational boundaries.
Preferred Skills
- Experience with AI infrastructure, inference-serving systems, or large-scale machine learning systems.
- Experience with compilers, runtimes, kernel optimization, or performance engineering.
- Familiarity with LLVM, MLIR, Triton, CUDA, or ROCm.
- Familiarity with GPUs, accelerators, hardware architecture, ISA concepts, or vendor toolchains.
- Experience with inference-serving frameworks or engines such as vLLM, SGLang, Triton Inference Server, or similar systems.
- Experience building developer platforms, external APIs, remote execution systems, or secure partner-facing infrastructure.
- Experience working with strategic cloud, hardware, or infrastructure partners.
Benefits
- Salary range of $293,000–$445,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and is committed to providing reasonable accommodations to applicants with disabilities. Background checks will be administered in accordance with applicable law.