Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Algorithms
Azure @ 4
Distributed Systems @ 4
GCP @ 4
Kubernetes @ 4
LLM @ 3
Machine Learning @ 4
Networking
Observability
Python @ 6
Rust @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Inference team builds and scales the systems that serve Claude to millions of users worldwide. The team operates compute-agnostic inference deployments across diverse AI accelerators and cloud platforms, covering intelligent request routing, fleet-wide orchestration, scaling, and networking.
Responsibilities
- Design, build, and maintain distributed systems serving Claude at global scale.
- Develop resilient systems that adapt to real-world events in real time.
- Build intelligent request routing, load balancing, and traffic management systems across thousands of accelerators and multiple cloud providers.
- Maximize compute efficiency and optimize fleet costs through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure for next-generation model development.
- Integrate new AI accelerator platforms and support inference for new model architectures.
- Design routing algorithms, manage multi-region deployments and geographic routing, and analyze observability data to tune production performance.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- A results-oriented approach, flexibility, and willingness to take on work outside the formal job description.
- Interest in learning about machine learning systems and infrastructure.
- Ability to work in environments where technical excellence drives business results and research breakthroughs.
- Experience with high-performance, large-scale distributed systems is preferred.
- Experience implementing and deploying machine learning systems at scale is preferred.
- Experience with load balancing, request routing, or traffic management systems is preferred.
- Familiarity with LLM inference optimization, batching, and caching strategies is preferred.
- Experience with Kubernetes and cloud infrastructure such as AWS, GCP, or Azure is preferred.
- Proficiency in Python or Rust is preferred.
- Bachelor's degree or equivalent combination of education, training, and experience in a relevant field, or equivalent demonstrated through coursework, training, or professional experience.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space. Staff are expected to work from an Anthropic office at least 25% of the time. Anthropic sponsors visas where possible and retains an immigration lawyer to assist with visa applications.
More jobs at Anthropic
Software Engineer, Business Technology
Anthropic · London, United Kingdom
GBP 255,000-325,000 per year
Repairs Program Lead - Data Center Operations
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Product Policy Manager, Product Risk
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 245,000-285,000 per year
Applied AI Engineer, Beneficial Deployments (Life Sciences)
Anthropic · New York City, United States, San Francisco, United States
USD 280,000-320,000 per year
Software Engineer, Business Technology
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Similar jobs
Staff + Senior Software Engineer, Inference Deployment
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Staff Software Engineer, Inference
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Staff + Senior Software Engineer, Inference
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Staff + Senior Software Engineer, Cloud Inference Launch Engineering
Anthropic · San Francisco, United States
USD 320,000-485,000 per year
Staff + Senior Software Engineer, Cloud Inference
Anthropic · San Francisco, United States
USD 320,000-485,000 per year