Staff+ Software Engineer, Safeguards ML Infrastructure

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AWS @ 4 Distributed Systems @ 7 GCP @ 4 LLM @ 3 Machine Learning Observability Python @ 1 Rust @ 1

Details

The Safeguards ML Infrastructure team designs, builds, and operates the production infrastructure powering Claude's safety systems. The team owns critical backend services on the token generation path and infrastructure used to configure these systems during model provisioning across first-party infrastructure, AWS Bedrock, GCP Vertex, and other platforms.

The role focuses on building and operating large-scale distributed systems, production platforms, tooling, and infrastructure used by other engineers. Familiarity with ML research or transformer architectures is not required.

Responsibilities

  • Design, build, and deploy backend services that are critical safety components on the token sampling and generation path.
  • Own and operate production serving infrastructure across multiple deployment platforms, including first-party infrastructure, AWS Bedrock, and GCP Vertex.
  • Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of Claude requests.
  • Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.
  • Reduce on-call and operational-duty toil by building automation, tooling, and self-service workflows that minimize manual operations.
  • Build and maintain a safety registry with full provenance, tracking what is running in production, on which model, and when and by whom it was deployed.
  • Implement automated post-deployment validation to ensure correctness across platforms.
  • Work closely with ML researchers to productionize new safety techniques and translate experimental work into reliable, scalable production systems.
  • Contribute to platform-agnostic deployment tooling that brings third-party platforms to parity with first-party operational maturity.

Requirements

  • Proficiency in Python; Rust experience is a plus.
  • Experience designing, building, and operating high-QPS systems at global scale.
  • Strong foundation in distributed systems, including replication, consistency tradeoffs, failure modes, and SLO management under load.
  • Meaningful on-call experience for production systems, including incident response and postmortem-driven improvements.
  • Hands-on experience deploying and operating AWS and GCP cloud platforms at scale.
  • Experience approaching infrastructure as a platform by building systems and abstractions that other engineers can use.
  • Minimum education of a bachelor's degree or equivalent combination of education, training, and experience.
  • A relevant field of study demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Eight or more years of industry software engineering experience.
  • Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls.
  • Demonstrated success reducing operational toil through automation, including transitioning teams from manual deployment processes to self-service pipelines.
  • Familiarity with LLM inference systems and the operational characteristics of transformer-based models.

Compensation

  • Annual salary: $320,000–$485,000 USD

Benefits

  • Hybrid work policy requiring staff to be in one of Anthropic's offices at least 25% of the time; some roles may require more office time.
  • Visa sponsorship is available, although sponsorship cannot be guaranteed for every role or candidate. Anthropic retains an immigration lawyer to assist with visas.
  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Office space for collaboration.

More jobs at Anthropic

Similar jobs