Software Engineer, GPU Infrastructure - HPC
SCRAPED
Used Tools & Technologies
Not specified
Required Skills & Competences ?
Go @ 5 Grafana @ 2 Linux @ 6 Prometheus @ 2 Python @ 5 SQL @ 3 Distributed Systems @ 3 Networking @ 3 ChatGPT @ 3 Pandas @ 3Details
About the team
The Fleet team at OpenAI supports the computing environment that powers research and product development. The team oversees large-scale systems spanning data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. The work supports internal research and external products like ChatGPT and prioritizes safety, reliability, and responsible AI deployment.
About the role
As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of OpenAI's compute fleet. Minimizing hardware failure is key to research training progress and stable services. This role involves troubleshooting state-of-the-art systems at scale, performing system-level investigations, and building automation for detection and remediation.
Responsibilities
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with clusters, networking, and infrastructure teams.
- Partner with external operators to ensure a high level of quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
Requirements
- Experience managing large-scale server environments.
- A balance of strengths in building and operationalizing systems.
- Proficiency in Python, Go, or similar languages.
- Strong Linux, networking, and server hardware knowledge.
- Comfort digging into noisy data with SQL, PromQL, and Pandas (or similar tools).
Prior hardware expertise is not required for this role.
Bonus skills
- Experience with low-level hardware details, protocols, and associated Linux tooling (e.g., PCIe, InfiniBand, networking, power management, kernel perf tuning).
- Knowledge of hardware management protocols (e.g., IPMI, Redfish).
- High-performance computing (HPC) or distributed systems experience.
- Prior experience developing, managing, or designing hardware.
- Familiarity with monitoring tools (e.g., Prometheus, Grafana).
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company emphasizes safety, diverse perspectives, and responsible deployment. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.
Benefits
- Base pay within the listed range, with total compensation including equity and performance-related bonuses for eligible employees.
- Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), paid medical and caregiver leave.
- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees.
- 13+ paid company holidays and multiple company office closures throughout the year.
- Mental health and wellness support; employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and meal delivery credits as eligible.
- Relocation support for eligible employees.
- Additional taxable fringe benefits (charitable donation matching, wellness stipends, etc.).