Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API
Debugging
Distributed Systems @ 3
GPU
Grafana @ 3
Kubernetes @ 3
Observability @ 3
Rust @ 6
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Storage Infrastructure team builds and operates the storage foundation behind OpenAI's demanding research and production workloads. The team works with research to design storage systems for rapidly evolving experiments and owns the platform end to end, including backend systems, user-facing services and APIs, and control planes that manage data placement, movement, and retention.
The stack spans cloud and in-house object stores across varied workload profiles, including GPU-attached systems and dedicated storage hardware. The team also builds a federation layer that unifies these backends behind a simple interface and routes each workload to the appropriate storage solution.
Responsibilities
- Build and operate storage services that underpin OpenAI's research infrastructure.
- Develop object storage systems across cloud and in-house environments.
- Build systems for cross-region data movement, replication, and recovery.
- Design lifecycle management capabilities that keep data durable, available, and cost-effective.
- Evolve the federation layer that unifies multiple backend systems behind a simple interface.
- Improve performance, reliability, and operational excellence across the platform.
- Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads.
- Own infrastructure end to end, including debugging and long-term reliability improvements.
Requirements
- Experience building or operating distributed systems in production.
- Experience with storage infrastructure, object stores, distributed filesystems, or other data-intensive backend systems.
- Strong production coding skills, ideally in Rust or another systems-oriented language.
- Comfort working with Kubernetes-based systems.
- Experience with tools such as Terraform, Grafana, or similar infrastructure and observability tooling.
Benefits
- Base salary of $230,000–$385,000 per year.
- Equity, performance-related bonuses for eligible employees, and benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and eligible meal delivery credits.
- Relocation support for eligible employees.
More jobs at OpenAI
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Android Systems Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 216,000-342,000 per year
Senior Staff Software Engineer, Identity
OpenAI · Mountain View, United States, San Francisco, United States
USD 345,000-405,000 per year
Analytics Engineer, GTM
OpenAI · San Francisco, United States, New York City, United States
USD 220,000-335,000 per year
Product Designer, Payments
OpenAI · San Francisco, United States
USD 245,000-310,000 per year
Similar jobs
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Software Engineer, Product Security Data Platforms
Stripe · Seattle, United States
USD 156,800-235,200 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
Senior Software Engineer - Postgres
ClickHouse · United States
USD 140,000-230,000 per year