Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Algorithms
Azure @ 4
Distributed Systems @ 4
GCP @ 4
Kubernetes @ 4
LLM @ 3
Machine Learning @ 4
Networking
Observability
Python @ 6
Rust @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Inference team builds and maintains the critical systems that serve Claude to millions of users worldwide. The team operates compute-agnostic inference deployments across diverse AI accelerators and cloud platforms, covering intelligent request routing, fleet-wide orchestration, scaling, networking, and model deployment.
The role focuses on maximizing compute efficiency while enabling high-performance inference infrastructure for research into next-generation models. Inference systems are performance-sensitive distributed systems serving hundreds of thousands of customers every day.
Responsibilities
- Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
- Develop resilient, flexible systems that adapt in real time to real-world events.
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
- Maximize compute efficiency across the fleet through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models to users.
- Provide high-performance inference infrastructure that enables researchers to develop next-generation models.
- Integrate new AI accelerator platforms and support inference for new model architectures.
Requirements
Minimum Qualifications
- Significant software engineering experience, particularly with distributed systems.
- A results-oriented approach, with a bias toward flexibility and impact.
- Willingness to take on work outside the formal job description when needed.
- Desire to learn more about machine learning systems and infrastructure.
- Ability to thrive in environments where technical excellence directly drives business results and research breakthroughs.
- Care about the societal impacts of the work.
Preferred Qualifications
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure, including AWS, GCP, or Azure.
- Proficiency in Python or Rust.
Education and Experience
- Bachelor's degree or an equivalent combination of education, training, and/or experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
- Required years of experience correlate with the internal job level requirements.
Representative Projects
- Designing intelligent routing algorithms that optimize request distribution across accelerators in different environments.
- Autoscaling the compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
- Building production-grade deployment pipelines for reliably releasing new models to millions of users.
- Contributing to new inference features.
- Supporting inference for new model architectures.
- Analyzing observability data to tune performance based on real-world production workloads.
- Managing multi-region deployments and geographic routing for global customers.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Anthropic expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas for eligible roles and candidates and retains an immigration lawyer to assist with the process.