Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Distributed Systems @ 3
GPU
Go @ 5
Grafana @ 2
HPC
InfiniBand @ 3
Linux @ 6
Networking @ 6
Pandas @ 3
Prometheus @ 2
Python @ 5
SQL @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Fleet team at OpenAI supports the computing environment that powers large-scale research and product development. The team oversees systems spanning data centers, GPUs, networking, and more, with a focus on availability, performance, efficiency, safety, reliability, and responsible AI deployment.
As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of OpenAI's compute fleet. You will investigate system-level issues, develop automated detection and remediation solutions, and maintain the health and efficiency of large-scale supercomputing infrastructure. Prior hardware expertise is not required.
Responsibilities
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with clusters, networking, and infrastructure teams.
- Partner with external operators to ensure a high level of quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
Requirements
- Experience managing large-scale server environments.
- A balance of strengths in building and operationalizing systems.
- Proficiency in Python, Go, or similar programming languages.
- Strong knowledge of Linux, networking, and server hardware.
- Comfort investigating noisy data with SQL, PromQL, Pandas, or other tools.
Bonus Skills
- Experience with low-level hardware details, protocols, and associated Linux tooling, including PCIe, InfiniBand, networking, power management, or kernel performance tuning.
- Knowledge of hardware management protocols such as IPMI and Redfish.
- High-performance computing or distributed systems experience.
- Experience developing, managing, or designing hardware.
- Familiarity with monitoring tools such as Prometheus and Grafana.
Benefits
- Medical, dental, and vision insurance for employees and families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, and office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.