Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Atlantis @ 7
Azure @ 6
BGP @ 4
Distributed Systems @ 6
GCP @ 6
Go @ 7
IaC
Kubernetes @ 6
Networking @ 4
Python @ 7
Rust @ 7
Security @ 4
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Infrastructure organization builds systems that support reliable AI research, safety experiments, and scaling Claude. The Cluster Infrastructure team owns the full lifecycle of compute clusters, including agent-driven provisioning and lifecycle management across major cloud providers and Anthropic datacenters. The systems provide high-bandwidth connectivity, secure-by-default configurations, and automated failure recovery.
Responsibilities
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management, including provisioning, updates, and decommissioning.
- Partner across teams to ensure new compute capacity is ingested on time.
- Align with partner teams on physical build-out and use cloud solutions to deliver high-bandwidth inter-cluster connectivity.
- Collaborate with security owners to ensure clusters are provisioned secure-by-default.
- Define and drive strategy for cluster scalability, homogeneity, and fault tolerance.
- Work with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy.
- Establish and evolve operational-excellence practices, including incident response, postmortem culture, and on-call health.
- Support the growth of engineers through technical mentorship and coaching.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, and AWS, GCP, or Azure.
- Strong proficiency in at least one systems language, such as Rust, Go, or Python.
- Infrastructure-as-code proficiency with Terraform.
- A track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively at all levels.
- 10+ years of software engineering experience preferred, including experience as a technical lead setting direction for a team.
- Experience operating large-scale compute infrastructure at hyperscale, including 100+ clusters and 10K+ nodes, preferred.
- Depth in one or more of Kubernetes internals, cluster provisioning and management systems, or cluster orchestration systems such as Mesos or Borg-like systems, preferred.
- Experience with cloud networking, including VPC design and peering, Shared VPC or Transit Gateway, Cloud Interconnect or Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing, and DDoS mitigation such as Cloud Armor or AWS Shield, preferred.
- Experience with cluster and host networking, including CNI such as Cilium, eBPF, NetworkPolicy, multi-NIC, sFlow, and service meshes such as Istio, Envoy, or Linkerd with mTLS, preferred.
- Experience with cluster security, including pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, and supply-chain or image provenance, preferred.
- Deep experience with infrastructure as code, including Terraform and Atlantis, and workflow orchestration such as Temporal and Argo Workflows, preferred.
- Skill in quickly understanding systems design tradeoffs and tracking rapidly evolving software systems.
- Bachelor's degree or an equivalent combination of education, training, and experience. The field of study should be relevant to the role through coursework, training, or professional experience.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
More jobs at Anthropic
AV Engineer
Anthropic · New York City, United States
USD 285,000-325,000 per year
IT Support Engineer
Anthropic · Dublin, Ireland
EUR 125,000-165,000 per year
AV Engineer
Anthropic · London, United Kingdom
GBP 210,000-250,000 per year
Forward Deployed Engineer
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 280,000-320,000 per year
Staff Software Engineer, Infrastructure (Distributed Systems)
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Similar jobs
Senior Staff+ Infrastructure Engineer, Cluster Infrastructure
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Spain, Ireland, Sweden, Germany, United Kingdom
GBP 91,000-1,114,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, Ireland, Sweden, United Kingdom
EUR 97,000-121,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Spain, Ireland, Sweden, Germany, United Kingdom
EUR 83,000-104,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year