Staff+ Site Reliability Engineer, Safeguards ML Infra

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AWS @ 4 Change Management @ 4 GCP @ 4 LLM @ 3 Machine Learning Python @ 6 Rust @ 1 SRE

Details

The Safeguards ML Infra team designs, builds, and operates the production infrastructure powering Claude's safety systems. The team owns critical backend services on the token generation path, deploys safeguards for new model launches, ships safety classifiers, verifies safeguards across 1P, AWS Bedrock, and GCP Vertex, and leads incident response.

The role focuses on production change management for safety-critical systems. Responsibilities include configuring and deploying safeguards for model releases, owning off-cycle safety classifier deployments, canarying changes, validating deployments, investigating discrepancies, managing rollback decisions, eliminating configuration drift, and automating launch and verification processes.

Responsibilities

  • Serve as launch captain for model releases by configuring and verifying safeguards and acting as the safeguards point of contact during release windows.
  • Own off-cycle deployments of new safety classifiers, including canary rollouts, post-deployment validation, and discrepancy investigation.
  • Verify that the correct safeguards are live on the correct models across 1P, AWS Bedrock, GCP Vertex, and other deployment platforms.
  • Automate launch runbooks, manual checks, and one-off deployments into tooling, continuous validation, and repeatable pipelines.
  • Build and maintain a safeguards registry with full provenance, including what is running, on which model and platform, and when and by whom it was deployed.
  • Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.
  • Use Claude to support safe agentic operations of safety-critical systems.

Requirements

  • Experience owning production change management at scale, including deployment pipelines, configuration management systems, and canary analysis.
  • Experience leading high-stakes releases as a launch captain, incident commander, or release owner.
  • Meaningful production on-call and incident-response experience, including postmortem-driven process and tooling improvements.
  • Hands-on experience deploying and operating cloud platforms at scale, particularly AWS and GCP.
  • Proficiency in Python.
  • A track record of reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
  • Minimum education of a bachelor's degree or equivalent combination of education, training, and experience.
  • Eight or more years of industry software engineering or site reliability engineering experience is a strong candidate qualification.
  • Experience with Rust is a plus.
  • Familiarity with LLM inference systems and transformer-based model operations is a plus; ML research or transformer architecture experience is not required.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration. The role requires staff to work from an office at least 25% of the time, with some roles requiring more office time.

More jobs at Anthropic

Similar jobs