Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS
Distributed Systems @ 4
GCP
Go @ 6
Java @ 6
Leadership @ 7
Machine Learning
Memcached @ 4
Observability @ 4
Python @ 6
Redis @ 4
Rust @ 6
Technical Leadership @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking experienced engineers to build its cache layer as a managed service from the ground up. The Caching team is part of the Databases organization and owns a managed Redis fleet, client libraries used across the company, and CDC-driven cache invalidation systems. The role will set the technical direction for caching across Anthropic, covering the data plane, developer experience, performance, reliability, and consistency.
Responsibilities
- Drive the technical direction for caching infrastructure used across Product and Research.
- Design, build, and operate a managed Redis fleet that scales to support millions of users across Claude's product ecosystem.
- Build client libraries and developer-facing abstractions that make correct caching the default for Anthropic engineers.
- Design and operate CDC-driven cache invalidation to keep cached data consistent with source-of-truth databases.
- Architect caching solutions operating across GCP, AWS, first-party deployments, and other environments.
- Optimize latency, hit rates, reliability, and cost efficiency on Anthropic's hottest paths.
- Build observability and tooling that makes cache behavior easy to understand and debug.
- Partner with product and research teams to understand access patterns and build infrastructure that accelerates their work.
- Make build-versus-buy decisions for caching technologies.
Requirements
- Significant experience building and operating production distributed systems as a software engineer.
- Deep knowledge of caching architectures, including invalidation strategies, consistency tradeoffs, and failure modes.
- Production experience operating Redis, Memcached, or similar in-memory data stores.
- Proficiency in at least one systems programming language, such as Go, Rust, Java, or C++, or Python at scale.
- A track record of leading large, complex infrastructure projects as an engineer or technical lead.
- Ability to balance rapid development with the reliability needs of production systems.
- Strong technical leadership and cross-functional collaboration skills.
- Preferred: 10+ years building and scaling distributed infrastructure, with 3+ years leading large-scale projects or teams.
- Preferred: Experience building managed infrastructure platforms or internal services consumed by many engineering teams.
- Preferred: Experience with change data capture, including Debezium or similar technologies, or streaming data infrastructure.
- Preferred: Experience operating Redis Cluster, Valkey, ElastiCache, Memorystore, or similar managed offerings at scale.
- Preferred: Experience designing client libraries or SDKs for internal infrastructure.
- Preferred: Experience scaling infrastructure during periods of rapid growth at high-growth companies.
- Preferred: Experience with multi-cloud or hybrid-cloud deployments.
- Preferred: Contributions to caching systems, database internals, or related open-source projects.
- Prior AI/ML infrastructure experience is not required.
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
Compensation
The annual salary range is $320,000–$485,000 USD.
Logistics
Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas for this role where possible and makes reasonable efforts to obtain a visa with the assistance of an immigration lawyer. Benefits include competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration.