Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Compliance
GPU @ 3
HPC @ 3
Linux @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. The team develops and maintains Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation across heterogeneous GPU fleets.
Responsibilities
- Build and maintain the Linux host software stack for large GPU clusters, including Ubuntu and OS images, kernel configuration and modules, drivers, packages, disks and storage configuration, and machine configuration.
- Design reproducible OS-image builds, package and repository workflows, and system configuration for heterogeneous bare-metal and cloud compute fleets.
- Integrate, test, and qualify kernels, modules, drivers, packages, and firmware across new and existing hardware platforms.
- Build safe canary, rollback, and recovery paths for system changes.
- Bring up new hardware platforms and compute SKUs, collaborating with hardware engineers and vendors to resolve firmware, driver, operating-system, and compatibility issues.
- Debug complex failures across boot and provisioning, firmware, disks, kernels and drivers, and workload interactions.
- Turn recurring failure modes into durable fixes, tests, and automation.
- Improve provisioning, repave, maintenance, and recovery correctness by eliminating manual host-by-host intervention.
- Build systems tooling and diagnostics to validate, troubleshoot, and operate the host software stack across the installed fleet.
Requirements
- Significant experience building, integrating, or operating production Linux systems, particularly Ubuntu, Debian, or another major Linux distribution.
- Deep experience in one or more of the following areas: Linux kernels, modules, or device drivers; Linux distribution, package, repository, or OS-image engineering; boot, provisioning, disks, or machine configuration; or firmware and driver integration, qualification, or rollout.
- Ability to write, debug, and maintain production-quality systems software and automation using appropriate systems languages, scripting, and Linux tooling.
- Understanding of how to build, test, package, qualify, and safely deliver system-level changes, including compatibility testing, staged rollout, rollback, and recovery.
- Ability to systematically debug failures across hardware, firmware, boot, operating-system, kernel, and driver boundaries.
- Effective collaboration with hardware and software partners.
Bonus Points
- Experience building bare-metal provisioning, machine-bootstrap, OS-reprovisioning, or system-lifecycle tooling.
- Experience with GPU systems, AI infrastructure, HPC clusters, accelerators, or other heterogeneous and performance-sensitive compute environments.
- Experience bringing up new server platforms or working with hardware and software vendors on firmware, driver, or operating-system compatibility issues.
- Contributions to Linux, kernel subsystems or modules, distributions, package ecosystems, drivers, firmware tooling, or other open-source systems software.
- Experience delivering or supporting system-software changes across large production compute fleets.
Benefits
- Base salary range of $230,000–$490,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax flexible spending and commuter accounts.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and paid office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
OpenAI is an equal opportunity employer committed to reasonable accommodations and compliance with applicable employment laws.
More jobs at OpenAI
Software Engineer, Monetization Data Systems
OpenAI · Mountain View, United States
USD 293,000-385,000 per year
Research Engineer / Research Scientist / AI Systems Engineer, RSI
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Researcher, Frontier Biological and Chemical Risks
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Research Engineer / Research Scientist - Personal AGI, Memory
OpenAI · San Francisco, United States
USD 295,000-555,000 per year
Research Engineer / Research Scientist - Personal AGI, Personalization
OpenAI · San Francisco, United States
USD 295,000-555,000 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Network Automation Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Data Center Manager
Nebius · Philadelphia, United States
USD 115,000-275,000 per year
CPU/Storage/PoP-WAN Program Manager
OpenAI · San Francisco, United States, Seattle, United States
USD 226,000-285,000 per year
Senior Technical Marketing Engineer - CAE Performance
Nvidia · Santa Clara, United States
USD 136,000-253,000 per year
Senior Storage Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
IT Infrastructure Engineer – RMA & Hardware Diagnostics
Nebius · Kansas City, United States
USD 112,700-140,800 per year