Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
AWS @ 4
Algorithms
Communication @ 7
Distributed Systems @ 3
GCP @ 4
Kubernetes @ 4
LLM @ 3
Machine Learning @ 4
Observability
Performance Optimization @ 3
Python @ 6
Rust @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s Inference team builds and maintains the critical systems that serve Claude to millions of users worldwide. The team operates compute-agnostic inference deployments and owns the stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.
The team focuses on maximizing compute efficiency while enabling high-performance inference infrastructure for AI research. The work involves complex distributed systems challenges across multiple accelerator families and cloud platforms.
As a Staff Software Engineer, you will work end to end to identify and address infrastructure blockers, support large-scale model serving, and enable AI research. Strong candidates will have familiarity with performance optimization, distributed systems, large-scale service orchestration, and intelligent request routing. Familiarity with LLM inference optimization, batching strategies, and multi-accelerator deployments is encouraged but not required.
Responsibilities
- Design intelligent routing algorithms to optimize request distribution across thousands of accelerators.
- Autoscale compute fleets to match supply and demand across production, research, and experimental workloads.
- Build production-grade deployment pipelines for releasing new models to millions of users.
- Integrate new AI accelerator platforms.
- Contribute to inference features such as structured sampling and prompt caching.
- Support inference for new model architectures.
- Analyze observability data to tune performance using real-world production workloads.
- Manage multi-region deployments and geographic routing for global customers.
- Address infrastructure blockers involved in serving Claude at scale and enabling AI research.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale is beneficial.
- Experience with load balancing, request routing, or traffic management systems is beneficial.
- Familiarity with LLM inference optimization, batching, and caching strategies is beneficial.
- Experience with Kubernetes and cloud infrastructure, including AWS or GCP, is beneficial.
- Proficiency with Python or Rust is beneficial.
- Bachelor’s degree or an equivalent combination of education, training, and experience.
- Education or professional experience in a field relevant to the role.
- Strong communication skills, flexibility, results orientation, and a bias toward impact.
- Interest in machine learning systems and infrastructure and the societal impacts of AI.
Benefits
- Annual salary of €295,000–€355,000 EUR.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
- Visa sponsorship may be available, with reasonable efforts and immigration lawyer support.
Applications are reviewed on a rolling basis, with no application deadline. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time.