Staff+ Site Reliability Engineer, Safeguards ML Infra

USD 405,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AWS @ 4 Change Management @ 4 GCP @ 4 LLM @ 3 Machine Learning Python @ 6 Rust @ 1 SRE

Details

The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. The team owns critical backend services on the token generation path and the operational work required to safely deploy safeguards for new model launches and new safety classifiers across 1P, AWS Bedrock, GCP Vertex, and other platforms.

This role focuses on configuring and deploying safeguards for model launches, owning off-cycle safety classifier deployments, canarying changes, verifying that safeguards are live on the correct models, authorizing rollbacks when necessary, and automating manual launch and verification processes into repeatable tooling and pipelines. Familiarity with ML research or transformer architectures is not required.

Responsibilities

  • Serve as launch captain for model releases by standing up, configuring, and verifying safeguards and acting as the safeguards point of contact during release windows.
  • Own off-cycle deployments of new safety classifiers, including canary rollouts, post-deployment validation, and discrepancy investigation.
  • Verify that the correct safeguards are live on the correct models across 1P, AWS Bedrock, GCP Vertex, and other deployment platforms.
  • Detect and eliminate configuration drift between deployment platforms.
  • Convert launch runbooks into tooling, manual checks into continuous validation, and one-off deployments into repeatable pipelines.
  • Build and maintain a safeguards registry with full provenance, including what is running in production, on which model and platform, and when and by whom it was deployed.
  • Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.
  • Help enable safe agentic operations for safety-critical systems.

Requirements

  • Experience owning production change management at scale, including deployment pipelines, configuration management systems, or canary analysis.
  • Experience running high-stakes releases as a launch captain, incident commander, or release owner.
  • Meaningful on-call experience for production systems, including incident response and postmortem-driven improvements.
  • A track record of converting postmortem action items into process and tooling improvements.
  • Hands-on experience deploying and operating cloud platforms at scale, including AWS and GCP.
  • Proficiency in Python.
  • A bachelor's degree or equivalent combination of education, training, and experience in a field relevant to the role.

Preferred Qualifications

  • Eight or more years of industry software engineering or site reliability engineering experience.
  • Demonstrated success reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
  • Experience running launch or production-readiness review processes across multiple teams.
  • Familiarity with LLM inference systems and the operational characteristics of transformer-based models.
  • Experience with Rust is a plus.

Logistics

  • Remote-friendly, with travel required.
  • Work locations: San Francisco, California; Seattle, Washington; or New York City, New York.
  • Employees are currently expected to work from an office at least 25% of the time, though some roles may require more office time.
  • Anthropic sponsors visas for this role when possible and makes reasonable efforts to provide visa support.

Compensation

  • Annual salary: $405,000–$485,000 USD.

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.

More jobs at Anthropic

Similar jobs