Staff Software Engineer, AI Reliability Engineering

📍 Dublin, Ireland
EUR 235,000-295,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 API Communication @ 6 Distributed Systems @ 7 GPU InfiniBand @ 4 LLM Machine Learning Networking @ 4 Observability @ 6 SRE @ 4

Details

Anthropic is seeking a Staff Software Engineer to join its AI Reliability Engineering (AIRE) team in Dublin. AIRE partners with teams across Anthropic to improve reliability across critical serving paths, including the SDK, network, API layers, serving infrastructure, and accelerators.

Responsibilities

  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving, which is critical for site reliability and Anthropic's safety commitments.

Requirements

  • Strong background in distributed systems, infrastructure, or reliability engineering.
  • Experience as a reliability-minded software engineer or site reliability engineer.
  • Ability to work effectively in unfamiliar systems during incidents and help drive resolution.
  • Holistic understanding of how systems compose and where system boundaries and seams occur.
  • Ability to build strong working relationships and collaborate across teams.
  • Excellent communication and collaboration skills.
  • Bachelor's degree or an equivalent combination of education, training, and/or experience.
  • Education or experience in a field relevant to the role.

Preferred Qualifications

  • Experience as an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure with more than 1,000 GPUs.
  • Experience with ML hardware accelerators such as GPUs, TPUs, or Trainium.
  • Understanding of ML-specific networking optimizations such as RDMA and InfiniBand.
  • Expertise in AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contributions to open-source infrastructure or ML tooling.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. The role follows a location-based hybrid policy, with staff expected to work from an Anthropic office at least 25% of the time. Anthropic sponsors visas where possible and retains an immigration lawyer to assist with the process.

More jobs at Anthropic

Similar jobs