Software Engineer, Fleet Hardware Health
SCRAPED
Used Tools & Technologies
Not specified
Required Skills & Competences ?
Go @ 5 Grafana @ 2 Linux @ 6 Prometheus @ 2 Python @ 5 SQL @ 3 Distributed Systems @ 3 Networking @ 3 Pandas @ 3Details
The Fleet team at OpenAI supports the computing environment that powers our research and product development. The team oversees large-scale systems spanning data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. This role focuses on the reliability and uptime of OpenAI's compute fleet, minimizing hardware failures and building automation for detection and remediation at scale.
Responsibilities
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with clusters, networking, and infrastructure teams.
- Partner with external operators to ensure a high level of quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
Requirements
- Experience managing large-scale server environments.
- A balance of strengths in building and operationalizing systems.
- Proficiency in Python, Go, or similar languages.
- Strong Linux, networking, and server hardware knowledge.
- Comfortable digging into noisy data with SQL, PromQL, and Pandas or other tools.
- Prior hardware expertise is not required for this role.
Bonus skills
- Experience with low-level hardware details, protocols, and Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning).
- Knowledge of hardware management protocols (e.g., IPMI, Redfish).
- High-performance computing (HPC) or distributed systems experience.
- Prior experience developing, managing, or designing hardware.
- Familiarity with monitoring tools (e.g., Prometheus, Grafana).
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We prioritize safety, reliability, and responsible AI deployment. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.
Background checks for applicants will be administered in accordance with applicable law. More information about OpenAI policies and applicant privacy is available to candidates through OpenAI links and policy documents provided during the application process.