Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 7
Distributed Systems @ 4
GPU @ 4
InfiniBand @ 4
Linux @ 4
Networking @ 4
Observability @ 6
Python @ 7
SQL @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Its mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling.
The team builds systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale.
Responsibilities
- Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
- Build and evolve health checks that detect, remediate, and verify failures at scale.
- Ensure critical health checks execute with minimal latency to maximize workload uptime.
- Investigate hardware failures and system-level issues across large-scale compute environments.
- Own node lifecycle workflows, including drain, quarantine, repair, RMA, and return-to-service processes.
- Build automation and tooling that enables global cluster management with minimal manual intervention.
- Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
Requirements
- 7+ years of industry experience in software or infrastructure engineering.
- Strong proficiency with Python and shell scripting.
- Experience building large-scale distributed systems or infrastructure platforms.
- Comfort analyzing noisy operational data using SQL, PromQL, or similar tooling.
- Experience building reproducible analyses and operational tooling.
- Strong systems debugging and operational instincts with an ownership mindset.
Bonus Qualifications
- Experience with low-level hardware systems and Linux tooling, including PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, or firmware/software debugging.
- Experience operating or debugging large-scale GPU or accelerator clusters.
- Expertise in network operations, observability, or systems telemetry.
- Experience with automated remediation systems or fleet lifecycle management.
- Experience improving reliability, utilization, or workload uptime in distributed compute environments.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics.
Background checks are administered in accordance with applicable law. OpenAI is committed to providing reasonable accommodations to applicants with disabilities.