Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Distributed Systems @ 4
Kubernetes @ 4
Memcached @ 6
Networking @ 4
Observability
Redis @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
At OpenAI, the Caching Infrastructure team builds a high-availability, multi-tenant caching layer that supports critical use cases across inference, identity, quota, and product experiences. The platform is designed to scale automatically with workload, minimize tail latency, and support diverse use cases.
The team is seeking an experienced engineer to help design and scale this infrastructure, with expertise in distributed caching systems, networking fundamentals, and Kubernetes-based service orchestration.
Responsibilities
- Design, build, and operate OpenAI's multi-tenant caching platform.
- Define the long-term vision and roadmap for caching as a core infrastructure capability, balancing performance, durability, and cost.
- Collaborate with infrastructure teams, including networking, observability, and databases, as well as product teams, to ensure the caching platform meets their needs.
Requirements
- 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems.
- Deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning.
- Production experience with Kubernetes, service meshes such as Envoy, and autoscaling systems.
- Ability to rigorously evaluate latency, reliability, throughput, and cost when designing platform capabilities.
- Ability to work in a fast-paced environment while balancing pragmatic engineering with long-term technical excellence.
Benefits
- Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.
More jobs at OpenAI
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Android Systems Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 216,000-342,000 per year
Senior Staff Software Engineer, Identity
OpenAI · Mountain View, United States, San Francisco, United States
USD 345,000-405,000 per year
Analytics Engineer, GTM
OpenAI · San Francisco, United States, New York City, United States
USD 220,000-335,000 per year
Product Designer, Payments
OpenAI · San Francisco, United States
USD 245,000-310,000 per year
Similar jobs
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Staff Software Engineer, Customer Administration
Coinbase · India
INR 9,424,500 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Staff+ Software Engineer, Caching
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Staff + Senior Software Engineer, Inference Deployment
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Software Engineer, Compute Infrastructure
OpenAI · New York City, United States, San Francisco, United States, Seattle, United States, London, United Kingdom
USD 230,000-405,000 per year
Staff+ Software Engineer, Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year