Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API
Communication @ 6
Distributed Systems @ 7
GPU
InfiniBand @ 4
LLM
Machine Learning
Networking @ 4
Observability @ 6
SRE @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Staff Software Engineer to join its AI Reliability Engineering (AIRE) team in Dublin. AIRE partners with teams across Anthropic to improve reliability across critical serving paths, including the SDK, network, API layers, serving infrastructure, and accelerators.
Responsibilities
- Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
- Design and implement monitoring and observability systems across the token path.
- Assist in designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.
- Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
- Support the reliability of safeguard model serving, which is critical for site reliability and Anthropic's safety commitments.
Requirements
- Strong background in distributed systems, infrastructure, or reliability engineering.
- Experience as a reliability-minded software engineer or site reliability engineer.
- Ability to work effectively in unfamiliar systems during incidents and help drive resolution.
- Holistic understanding of how systems compose and where system boundaries and seams occur.
- Ability to build strong working relationships and collaborate across teams.
- Excellent communication and collaboration skills.
- Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Education or experience in a field relevant to the role.
Preferred Qualifications
- Experience as an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
- Experience operating large-scale model serving or training infrastructure with more than 1,000 GPUs.
- Experience with ML hardware accelerators such as GPUs, TPUs, or Trainium.
- Understanding of ML-specific networking optimizations such as RDMA and InfiniBand.
- Expertise in AI-specific observability tools and frameworks.
- Experience with chaos engineering and systematic resilience testing.
- Contributions to open-source infrastructure or ML tooling.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. The role follows a location-based hybrid policy, with staff expected to work from an Anthropic office at least 25% of the time. Anthropic sponsors visas where possible and retains an immigration lawyer to assist with the process.
More jobs at Anthropic
Developer Relations
Anthropic · San Francisco, United States, New York City, United States
USD 290,000-435,000 per year
AV Operations Specialist
Anthropic · San Francisco, United States, New York City, United States
USD 230,000-285,000 per year
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Staff Software Engineer, Claude Code
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-625,000 per year
Applied AI Architect, Enterprise Tech
Anthropic · San Francisco, United States, New York City, United States
USD 240,000-315,000 per year
Similar jobs
Staff Software Engineer, AI Reliability
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 325,000-485,000 per year
Staff Software Engineer, AI Reliability Engineering
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior AI Infrastructure Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Senior Engineering Manager, Object Storage - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year