Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 1
Azure @ 3
CI/CD
GPU
Kubernetes @ 3
Machine Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
This role supports OpenAI's fleet infrastructure team, which operates a large-scale GPU fleet for general-purpose model training and deployment. The team builds scheduling and quota systems, Kubernetes cluster provisioning and upgrade automation, service frameworks, deployment systems, and high-performance snapshot delivery from blob storage to hardware caching.
Responsibilities
- Design, implement, deploy, and operate infrastructure systems for model deployment and training.
- Design and operate compute fleet components, including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
- Interface with researchers and product teams to understand workload requirements.
- Collaborate with hardware, infrastructure, and business teams to provide highly utilized and reliable services.
- Build user-friendly scheduling and quota systems to maximize GPU utilization.
- Develop push-button automation for Kubernetes cluster provisioning and upgrades.
- Support research workflows with service frameworks and deployment systems.
- Improve model startup times through high-performance snapshot delivery across blob storage and hardware caching.
Requirements
- Experience with hyperscale compute systems.
- Strong programming skills.
- Experience working with public clouds, especially Azure.
- Experience working with Kubernetes.
- An execution-focused mentality combined with rigorous attention to user requirements.
- An understanding of AI/ML workloads is a bonus.
Work Arrangement
- Based in San Francisco, California.
- Hybrid work model with three days per week in the office.
- Relocation assistance is available to eligible new employees.
Benefits
- Base salary range of $230,000–$490,000 per year, plus equity.
- Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and paid office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Strategic Delivery Lead, Intelligence Community
OpenAI · Washington, United States
USD 266,000-370,000 per year
Product Manager, Youth
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Software Engineer, API Safety
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Head of Marketplace
OpenAI · New York City, United States, San Francisco, United States
USD 400,000-445,000 per year
Research Engineer / Research Scientist, Health
OpenAI · San Francisco, United States
USD 295,000-555,000 per year
Similar jobs
Technical Program Manager, Infrastructure
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior DevOps Engineer, Platform Engineering
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Infrastructure Security Engineer
SpaceXAI · Washington, United States, Austin, United States, New York City, United States, Palo Alto, United States
USD 100,000-258,000 per year