Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
ChatGPT
Codex
Distributed Systems
FastAPI
Kubernetes
Machine Learning @ 6
Terraform @ 3
gRPC
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Agent Infrastructure team builds systems for training and deploying highly useful AI agents. This includes environments where AI models can execute code, debug issues, and develop software, as well as the production platform powering products such as Codex, Operator, tool use in ChatGPT, and future agentic products.
The role involves working closely with research and product teams to build and scale infrastructure for training complex agentic models and launching agents to hundreds of millions of users. Responsibilities include developing new infrastructure and integrations, scaling capabilities to large compute clusters, and maintaining the production platform on which agents run.
Responsibilities
- Contribute to an in-house container orchestration platform designed to scale beyond systems such as Kubernetes.
- Develop and maintain FastAPI and gRPC APIs for agentic infrastructure used in training and production.
- Use Terraform to provision and evolve complex research and production infrastructure.
- Collaborate with research teams to establish and optimize systems for novel AI training runs and experimental applications.
- Identify bottlenecks and engineer performance improvements for large-scale machine learning infrastructure.
- Build systems from 0-to-1 and scale them to large production environments.
- Optimize globally distributed systems, virtualization efficiency, and runtime performance.
Requirements
- Deep experience working on large-scale machine learning infrastructure.
- Experience reasoning about training at scale and optimizing system performance in training environments.
- Experience building and scaling systems for complex, ambiguous technical problems.
- Knowledge of cloud platforms and infrastructure-as-code technologies such as Terraform.
- Deep technical expertise in virtualization and containerization technologies, such as Kata, Firecracker, gVisor, or Sysbox.
- Passion for optimizing runtime performance.
Benefits
- Base salary of $230,000–$385,000 per year.
- Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, parking, and transit expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and paid office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
More jobs at OpenAI
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Android Systems Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 216,000-342,000 per year
Senior Staff Software Engineer, Identity
OpenAI · Mountain View, United States, San Francisco, United States
USD 345,000-405,000 per year
Analytics Engineer, GTM
OpenAI · San Francisco, United States, New York City, United States
USD 220,000-335,000 per year
Product Designer, Payments
OpenAI · San Francisco, United States
USD 245,000-310,000 per year
Similar jobs
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior System Software Engineer for Cloud – GeForce NOW Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Staff Software Engineer, Gov
OpenAI · Washington, United States, San Francisco, United States, Seattle, United States
USD 207,000-385,000 per year
Software Engineer, Developer Productivity
OpenAI · San Francisco, United States, New York City, United States, Seattle, United States
USD 210,000-490,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Engineer, Ads Business Manager
Reddit · United States
USD 217,000-303,900 per year