Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 3
Debugging @ 3
Distributed Systems @ 3
LLM @ 3
Observability @ 3
Profiling @ 3
Python @ 6
Rust @ 6
SGLang
System Architecture
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for advanced AI workloads. The team is developing future generations of AI-native silicon and tightly integrated systems to power frontier models and OpenAI’s supercomputing platform.
You will build the model runtime within the inference engine that executes complex frontier models at scale on OpenAI’s custom silicon. The runtime will connect models running on the hardware with the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing throughput, latency, utilization, and reliability.
You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, customized and optimized for OpenAI’s AI accelerator.
Responsibilities
- Design and implement the LLM inference runtime for frontier models running on custom silicon.
- Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
- Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
- Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
- Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.
- Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
- Create profiling, observability, benchmarking, and performance-modeling tools that make runtime behavior measurable and actionable.
- Debug complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
- Turn workload insights into clear requirements for future generations of silicon and system architecture.
Requirements
- Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- Experience building or optimizing runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
- Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- Ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
- Experience profiling and debugging performance across multiple layers of a hardware-software stack.
- Ability to design clean abstractions while retaining the low-level control needed to extract performance from specialized hardware.
- Ability to work across model, systems, compiler, kernel, and hardware teams to resolve ambiguous technical problems.
- Commitment to production quality, including correctness, observability, reliability, maintainability, and graceful behavior at scale.
- Candidates may need to meet certain legal status requirements under U.S. export control laws and regulations.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.
Benefits
- Base salary range of $266,000–$445,000 per year, plus equity and performance-related bonuses for eligible employees.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses, commuter expenses, and related benefits.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may be provided, including charitable donation matching and wellness stipends.