Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Change Management @ 4
GCP @ 4
LLM @ 3
Machine Learning
Python @ 6
Rust @ 1
SRE
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Safeguards ML Infra team designs, builds, and operates the production infrastructure powering Claude's safety systems. The team owns critical backend services on the token generation path, deploys safeguards for new model launches, ships safety classifiers, verifies safeguards across 1P, AWS Bedrock, and GCP Vertex, and leads incident response.
The role focuses on production change management for safety-critical systems. Responsibilities include configuring and deploying safeguards for model releases, owning off-cycle safety classifier deployments, canarying changes, validating deployments, investigating discrepancies, managing rollback decisions, eliminating configuration drift, and automating launch and verification processes.
Responsibilities
- Serve as launch captain for model releases by configuring and verifying safeguards and acting as the safeguards point of contact during release windows.
- Own off-cycle deployments of new safety classifiers, including canary rollouts, post-deployment validation, and discrepancy investigation.
- Verify that the correct safeguards are live on the correct models across 1P, AWS Bedrock, GCP Vertex, and other deployment platforms.
- Automate launch runbooks, manual checks, and one-off deployments into tooling, continuous validation, and repeatable pipelines.
- Build and maintain a safeguards registry with full provenance, including what is running, on which model and platform, and when and by whom it was deployed.
- Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.
- Use Claude to support safe agentic operations of safety-critical systems.
Requirements
- Experience owning production change management at scale, including deployment pipelines, configuration management systems, and canary analysis.
- Experience leading high-stakes releases as a launch captain, incident commander, or release owner.
- Meaningful production on-call and incident-response experience, including postmortem-driven process and tooling improvements.
- Hands-on experience deploying and operating cloud platforms at scale, particularly AWS and GCP.
- Proficiency in Python.
- A track record of reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
- Minimum education of a bachelor's degree or equivalent combination of education, training, and experience.
- Eight or more years of industry software engineering or site reliability engineering experience is a strong candidate qualification.
- Experience with Rust is a plus.
- Familiarity with LLM inference systems and transformer-based model operations is a plus; ML research or transformer architecture experience is not required.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration. The role requires staff to work from an office at least 25% of the time, with some roles requiring more office time.