Senior Staff+ Infrastructure Engineer, Cluster Infrastructure
at Anthropic
USD 405,000-485,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Atlantis @ 7
Azure @ 6
BGP @ 4
Distributed Systems @ 6
GCP @ 6
Go @ 7
IaC
Kubernetes @ 6
Networking @ 4
Python @ 7
Rust @ 7
Security @ 4
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Infrastructure organization builds systems that support reliable AI research, safety experiments, and scaling Claude. The Cluster Infrastructure team owns the full lifecycle of compute clusters, including agent-driven provisioning and lifecycle management across major cloud providers and Anthropic datacenters. These systems provide high-bandwidth connectivity, secure-by-default configurations, and automated failure recovery.
Responsibilities
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management, including provisioning, updates, and decommissioning.
- Partner across teams to ensure new compute capacity is ingested on time.
- Align with partner teams on physical build-out and use cloud solutions to provide high-bandwidth inter-cluster connectivity.
- Collaborate with security owners to ensure clusters are provisioned secure-by-default.
- Define and drive strategy for cluster scalability, homogeneity, and fault tolerance.
- Work with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy.
- Establish and evolve operational-excellence practices, including incident response, postmortem culture, and on-call health.
- Support the growth of engineers through technical mentorship and coaching.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, AWS, GCP, or Azure.
- Strong proficiency in at least one systems language, such as Rust, Go, or Python.
- Proficiency with Terraform and infrastructure as code.
- Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively at all levels.
- Bachelor’s degree or an equivalent combination of education, training, and/or experience in a relevant field.
Preferred Qualifications
- 12+ years of software engineering experience, including time as a technical lead setting direction for a team.
- Experience operating large-scale compute infrastructure at hyperscale, including 100+ clusters and 10K+ nodes.
- Expertise in Kubernetes internals, cluster provisioning and management systems, or cluster orchestration systems such as Mesos or Borg-like systems.
- Experience with cloud networking, including VPC design and peering, Shared VPC/Transit Gateway, Cloud Interconnect/Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing, and DDoS mitigation such as Cloud Armor or AWS Shield.
- Experience with cluster and host networking, including CNI such as Cilium, eBPF, NetworkPolicy, multi-NIC, sFlow, service mesh technologies such as Istio, Envoy, or Linkerd, and mTLS.
- Experience with cluster security, including pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, and supply-chain/image provenance.
- Deep experience with infrastructure as code, including Terraform and Atlantis.
- Experience with workflow orchestration, including Temporal and Argo Workflows.
- Skill in quickly understanding systems design tradeoffs and tracking rapidly evolving software systems.
Compensation
The annual salary range is $405,000–$485,000 USD.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration.
Logistics
- The role follows a location-based hybrid policy. Staff are expected to be in one of Anthropic’s offices at least 25% of the time, although some roles may require more office time.
- Anthropic explicitly sponsors visas for eligible roles and candidates and retains an immigration lawyer to provide support.
More jobs at Anthropic
Staff Software Engineer, Claude Code
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-625,000 per year
Applied AI Architect, Enterprise Tech
Anthropic · San Francisco, United States, New York City, United States
USD 240,000-315,000 per year
Finance Systems Engineer, Tax
Anthropic · San Francisco, United States, Seattle, United States
USD 205,000-270,000 per year
AI Fluency Education Lead
Anthropic · San Francisco, United States, New York City, United States
USD 270,000-365,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Similar jobs
Staff+ Infrastructure Engineer, Cluster Infrastructure
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Member of Technical Staff (Software Engineer, Cloud Infrastructure)
Perplexity AI · Palo Alto, United States, San Francisco, United States, New York City, United States
USD 220,000-405,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Security Engineer, Application & Platform Security
Sentry · San Francisco, United States
USD 190,000-280,000 per year