Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Algorithms
CI/CD @ 3
ChatGPT
Chef @ 3
GPU
Kubernetes @ 3
LLM
Linux @ 3
Networking
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Fleet team at OpenAI supports the computing environment that powers cutting-edge research and product development. The team oversees large-scale systems spanning data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. This work enables OpenAI's models to operate at scale, supporting internal research and products such as ChatGPT, with a focus on safety, reliability, and responsible AI deployment.
About the Role
The Software Engineer, Operating Systems & Orchestration will build systems to manage hardware, configurations, vendors, and the people interacting with infrastructure. The role involves designing and developing solutions that integrate individual nodes and servers into unified clusters, helping advance AI research by streamlining the overall research user experience.
This role is based in San Francisco, California. OpenAI uses a hybrid work model requiring three days in the office per week and offers relocation assistance to new employees.
Responsibilities
- Design and build systems to manage cloud and bare-metal fleets at scale.
- Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms.
- Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows.
- Automate infrastructure processes to reduce repetitive toil and improve system reliability.
- Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack.
- Continuously improve tools, automation, processes, and documentation to enhance operational efficiency.
Requirements
- Strong software engineering skills with experience in large-scale infrastructure environments.
- Broad knowledge of cluster-level systems, such as Kubernetes, CI/CD pipelines, Terraform, and cloud providers.
- Deep expertise in server-level systems, including systems, containerization, Chef, Linux kernels, firmware management, and host routing.
- Passion for optimizing the performance and reliability of large compute fleets.
- Ability to thrive in dynamic environments and solve complex infrastructure challenges.
- Commitment to automation, efficiency, and continuous improvement.
Benefits
- Base salary range of $230,000–$490,000 per year.
- Equity, performance-related bonuses for eligible employees, and additional benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax FSA, dependent care FSA, commuter, parking, and transit accounts.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.