Staff+ Infrastructure Engineer, Cluster Infrastructure

GBP 325,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI AWS @ 4 Atlantis @ 7 Azure @ 6 BGP @ 4 Distributed Systems @ 6 GCP @ 6 Go @ 7 IaC Kubernetes @ 6 Networking @ 4 Python @ 7 Rust @ 7 Security @ 4 Terraform @ 6

Details

Anthropic's Infrastructure organization builds systems that support reliable AI research, safety experiments, and scaling Claude. The Cluster Infrastructure team owns the full lifecycle of compute clusters, including agent-driven provisioning and lifecycle management across major cloud providers and Anthropic datacenters. The systems provide high-bandwidth connectivity, secure-by-default configurations, and automated failure recovery.

Responsibilities

  • Own the technical strategy and roadmap for agent-driven cluster lifecycle management, including provisioning, updates, and decommissioning.
  • Partner across teams to ensure new compute capacity is ingested on time.
  • Align with partner teams on physical build-out and use cloud solutions to deliver high-bandwidth inter-cluster connectivity.
  • Collaborate with security owners to ensure clusters are provisioned secure-by-default.
  • Define and drive strategy for cluster scalability, homogeneity, and fault tolerance.
  • Work with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy.
  • Establish and evolve operational-excellence practices, including incident response, postmortem culture, and on-call health.
  • Support the growth of engineers through technical mentorship and coaching.

Requirements

  • Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, and AWS, GCP, or Azure.
  • Strong proficiency in at least one systems language, such as Rust, Go, or Python.
  • Infrastructure-as-code proficiency with Terraform.
  • A track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels.
  • 10+ years of software engineering experience preferred, including experience as a technical lead setting direction for a team.
  • Experience operating large-scale compute infrastructure at hyperscale, including 100+ clusters and 10K+ nodes, preferred.
  • Depth in one or more of Kubernetes internals, cluster provisioning and management systems, or cluster orchestration systems such as Mesos or Borg-like systems, preferred.
  • Experience with cloud networking, including VPC design and peering, Shared VPC or Transit Gateway, Cloud Interconnect or Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing, and DDoS mitigation such as Cloud Armor or AWS Shield, preferred.
  • Experience with cluster and host networking, including CNI such as Cilium, eBPF, NetworkPolicy, multi-NIC, sFlow, and service meshes such as Istio, Envoy, or Linkerd with mTLS, preferred.
  • Experience with cluster security, including pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, and supply-chain or image provenance, preferred.
  • Deep experience with infrastructure as code, including Terraform and Atlantis, and workflow orchestration such as Temporal and Argo Workflows, preferred.
  • Skill in quickly understanding systems design tradeoffs and tracking rapidly evolving software systems.
  • Bachelor's degree or an equivalent combination of education, training, and experience. The field of study should be relevant to the role through coursework, training, or professional experience.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs