Staff Software Engineer, AI Reliability Engineering

GBP 325,000-390,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 API Communication @ 7 Distributed Systems @ 7 GPU InfiniBand @ 4 LLM Machine Learning Networking @ 4 Observability @ 6 SRE @ 4

Details

Anthropic’s AI Reliability Engineering (AIRE) team partners with teams across the company to improve reliability across critical Claude serving paths, including the SDK, network, API layers, serving infrastructure, and accelerators. The team works alongside partner teams during incidents and on projects to make systems more robust and resilient.

Responsibilities

  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving, which is critical for site reliability and Anthropic’s safety commitments.

Requirements

  • Strong background in distributed systems, infrastructure, or reliability engineering.
  • Reliability-minded software engineering or Site Reliability Engineering experience.
  • Ability to work comfortably with unfamiliar systems during incidents and help drive resolution.
  • Holistic understanding of how systems compose and where system seams occur.
  • Ability to build lasting relationships and collaborate across teams.
  • Strong communication and collaboration skills.
  • A bachelor’s degree or equivalent combination of education, training, and experience.
  • Education or experience in a field relevant to the role.

Preferred Qualifications

  • Experience as an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure with more than 1,000 GPUs.
  • Experience with ML hardware accelerators, including GPUs, TPUs, or Trainium.
  • Understanding of ML-specific networking optimizations such as RDMA and InfiniBand.
  • Expertise with AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contributions to open-source infrastructure or ML tooling.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.

More jobs at Anthropic

Similar jobs