Staff+ Software Engineer, Infrastructure (Distributed Systems)
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Communication @ 7
Data Pipelines
Distributed Systems @ 4
GCP @ 4
GPU
Go @ 7
Java @ 7
Kubernetes @ 4
Linux @ 4
Machine Learning @ 4
NCCL @ 4
Networking @ 4
Observability
Python @ 7
Rust @ 7
Scoping @ 6
Security @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Infrastructure organization builds and operates the distributed systems that train, serve, and secure its AI models. This includes data pipelines, large-scale Kubernetes clusters, databases, observability systems, and developer tooling used across the company. The role involves independently scoping and leading complex infrastructure projects, making architectural decisions, and partnering with research and product teams. Team placement is determined after the interview process based on the candidate's interests and experience and organizational needs.
Responsibilities
- Independently scope and lead complex, multi-month infrastructure projects from an ambiguous starting point through production.
- Make architectural decisions that shape the infrastructure foundation used by other engineers and teams.
- Drive alignment on technical direction across multiple teams and ambiguous problem spaces.
- Partner with research and product teams to understand infrastructure and compute needs and translate them into technical designs.
- Own the reliability, scalability, and security of systems as usage and complexity grow.
- Set technical strategy and infrastructure standards for the team.
- Build and improve operational processes, including incident response, postmortems, and on-call rotations.
- Mentor engineers and help raise the team's technical bar.
Requirements
- Experience designing, building, and operating large-scale distributed systems or infrastructure in production.
- Track record of independently scoping and delivering complex, ambiguous, multi-month technical projects.
- Experience making architectural decisions that other engineers and teams build upon.
- Strong software engineering fundamentals and proficiency in at least one programming language, such as Python, Rust, Go, or Java.
- Experience with modern cloud infrastructure, including Kubernetes and infrastructure-as-code, on AWS and/or GCP.
- Strong written and verbal communication skills, including experience driving alignment across multiple teams or stakeholders.
- A bachelor's degree or equivalent combination of education, training, and experience. The field of study should be relevant to the role as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- 10+ years of software engineering experience, excluding internships.
- Experience with machine learning infrastructure, including GPUs, TPUs, or Trainium, and associated networking infrastructure such as NCCL.
- Low-level systems experience, such as Linux kernel tuning or eBPF.
- Background in security or privacy engineering best practices.
- Prior experience as a technical lead or mentor for other engineers.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration.
Additional Information
- Applications are reviewed on a rolling basis; there is no application deadline.
- The location-based hybrid policy currently expects staff to work from one of the company's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and states that it makes every reasonable effort to obtain a visa for candidates who receive an offer, although sponsorship is not successful for every role and candidate.
- Anthropic is a public benefit corporation headquartered in San Francisco.