Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Distributed Systems @ 7
GCP @ 4
LLM @ 3
Machine Learning
Observability
Python @ 1
Rust @ 1
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Safeguards ML Infrastructure team designs, builds, and operates the production infrastructure powering Claude's safety systems. The team owns critical backend services on the token generation path and infrastructure used to configure these systems during model provisioning across first-party infrastructure, AWS Bedrock, GCP Vertex, and other platforms.
The role focuses on building and operating large-scale distributed systems, production platforms, tooling, and infrastructure used by other engineers. Familiarity with ML research or transformer architectures is not required.
Responsibilities
- Design, build, and deploy backend services that are critical safety components on the token sampling and generation path.
- Own and operate production serving infrastructure across multiple deployment platforms, including first-party infrastructure, AWS Bedrock, and GCP Vertex.
- Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of Claude requests.
- Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.
- Reduce on-call and operational-duty toil by building automation, tooling, and self-service workflows that minimize manual operations.
- Build and maintain a safety registry with full provenance, tracking what is running in production, on which model, and when and by whom it was deployed.
- Implement automated post-deployment validation to ensure correctness across platforms.
- Work closely with ML researchers to productionize new safety techniques and translate experimental work into reliable, scalable production systems.
- Contribute to platform-agnostic deployment tooling that brings third-party platforms to parity with first-party operational maturity.
Requirements
- Proficiency in Python; Rust experience is a plus.
- Experience designing, building, and operating high-QPS systems at global scale.
- Strong foundation in distributed systems, including replication, consistency tradeoffs, failure modes, and SLO management under load.
- Meaningful on-call experience for production systems, including incident response and postmortem-driven improvements.
- Hands-on experience deploying and operating AWS and GCP cloud platforms at scale.
- Experience approaching infrastructure as a platform by building systems and abstractions that other engineers can use.
- Minimum education of a bachelor's degree or equivalent combination of education, training, and experience.
- A relevant field of study demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Eight or more years of industry software engineering experience.
- Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls.
- Demonstrated success reducing operational toil through automation, including transitioning teams from manual deployment processes to self-service pipelines.
- Familiarity with LLM inference systems and the operational characteristics of transformer-based models.
Compensation
- Annual salary: $320,000–$485,000 USD
Benefits
- Hybrid work policy requiring staff to be in one of Anthropic's offices at least 25% of the time; some roles may require more office time.
- Visa sponsorship is available, although sponsorship cannot be guaranteed for every role or candidate. Anthropic retains an immigration lawyer to assist with visas.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
More jobs at Anthropic
Developer Relations
Anthropic · San Francisco, United States, New York City, United States
USD 290,000-435,000 per year
AV Operations Specialist
Anthropic · San Francisco, United States, New York City, United States
USD 230,000-285,000 per year
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Staff Software Engineer, Claude Code
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-625,000 per year
Applied AI Architect, Enterprise Tech
Anthropic · San Francisco, United States, New York City, United States
USD 240,000-315,000 per year
Similar jobs
Staff + Senior Software Engineer, Inference Deployment
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Staff + Senior Software Engineer, Cloud Inference Launch Engineering
Anthropic · San Francisco, United States
USD 320,000-485,000 per year
Staff + Senior Software Engineer, Cloud Inference
Anthropic · San Francisco, United States
USD 320,000-485,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Software Engineer, Inference
Anthropic · Dublin, Ireland
EUR 295,000-355,000 per year
Staff Software Engineer, Inference
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year