Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Algorithms
Azure @ 4
Distributed Systems @ 4
GCP @ 4
GPU
Kubernetes @ 4
LLM @ 3
Machine Learning @ 4
Observability
Python @ 6
Rust @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Inference team builds and maintains the critical systems that serve Claude to millions of users worldwide. The team operates large-scale, compute-agnostic inference deployments spanning intelligent request routing, fleet-wide orchestration, diverse AI accelerators, and multiple cloud platforms.
Inference systems are highly performance-sensitive distributed systems serving hundreds of thousands of customers every day. The role involves solving complex distributed systems challenges across multiple accelerator families and emerging AI hardware.
Responsibilities
- Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
- Develop resilient and flexible systems that adapt in real time to real-world events.
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
- Maximize compute efficiency across the fleet through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models to users.
- Provide high-performance inference infrastructure that enables researchers to develop next-generation models.
- Integrate new AI accelerator platforms and support inference for new model architectures.
- Design routing algorithms that optimize request distribution across accelerators and environments.
- Autoscale the compute fleet to dynamically match supply with demand.
- Analyze observability data to tune performance based on real-world production workloads.
- Manage multi-region deployments and geographic routing for global customers.
Requirements
- Proficiency in Python or Rust.
- Software engineering experience building and operating distributed systems in production.
- Working knowledge of containerized infrastructure, such as Kubernetes.
- Experience with at least one major cloud platform: AWS, GCP, or Azure.
- Results-oriented approach with flexibility and a focus on impact.
- Willingness to take on responsibilities beyond the formal job description.
- Desire to learn more about machine learning systems and infrastructure.
- Significant experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience building load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Deep experience operating Kubernetes and cloud infrastructure at scale.
- Experience with AI accelerator platforms, including GPUs, TPUs, or emerging hardware.
- A bachelor's degree or equivalent combination of education, training, and experience in a field relevant to the role, as demonstrated through coursework, training, or professional experience.
- Minimum years of experience correlate with the internal job-level requirements.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration.
Additional Information
Applications are reviewed on a rolling basis with no stated deadline. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas and will make reasonable efforts to support visa applications, although sponsorship is not guaranteed for every role or candidate.