Software Engineer, Fleet Hardware Health
Used Tools & Technologies
Not specified
Required Skills & Competences
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 β basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 β daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 β you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 β exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Go @ 5
Grafana @ 2
Linux @ 6
Prometheus @ 2
Python @ 5
SQL @ 3
Distributed Systems @ 3
Networking @ 3
Pandas @ 3
- 1-2 β basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 β daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 β you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 β exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Fleet team at OpenAI supports the computing environment that powers our research and product development. The team oversees large-scale systems spanning data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. This role focuses on the reliability and uptime of OpenAI's compute fleet, minimizing hardware failures and building automation for detection and remediation at scale.
Responsibilities
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with clusters, networking, and infrastructure teams.
- Partner with external operators to ensure a high level of quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
Requirements
- Experience managing large-scale server environments.
- A balance of strengths in building and operationalizing systems.
- Proficiency in Python, Go, or similar languages.
- Strong Linux, networking, and server hardware knowledge.
- Comfortable digging into noisy data with SQL, PromQL, and Pandas or other tools.
- Prior hardware expertise is not required for this role.
Bonus skills
- Experience with low-level hardware details, protocols, and Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning).
- Knowledge of hardware management protocols (e.g., IPMI, Redfish).
- High-performance computing (HPC) or distributed systems experience.
- Prior experience developing, managing, or designing hardware.
- Familiarity with monitoring tools (e.g., Prometheus, Grafana).
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We prioritize safety, reliability, and responsible AI deployment. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.
Background checks for applicants will be administered in accordance with applicable law. More information about OpenAI policies and applicant privacy is available to candidates through OpenAI links and policy documents provided during the application process.