Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 6
AWS @ 4
Communication @ 7
Consul @ 4
Distributed Systems @ 4
GCP @ 4
GPU
Go @ 6
IaC
Kubernetes @ 7
Linux @ 4
Machine Learning
NCCL @ 3
Networking @ 3
Python @ 6
Rust @ 6
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic operates one of the industry's largest AI compute fleets across multiple cloud providers and datacenters. The Kubernetes Platform team owns the control plane supporting these fleets, including the scheduler, control-plane components, and core cluster services.
The role focuses on building reliable, correct, and highly available infrastructure capable of supporting frontier AI model training, research, and serving at large scale.
Responsibilities
- Own, operate, and extend the Kubernetes scheduler for accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption.
- Scale the Kubernetes control plane, including the apiserver, etcd, and controller-manager, to support clusters far beyond typical limits.
- Identify and address performance bottlenecks in large-scale control-plane systems.
- Design, build, and operate core cluster services such as service discovery.
- Build and maintain custom controllers, operators, and CRDs.
- Partner with research, training, and inference teams to understand workload requirements and translate them into platform capabilities.
- Collaborate with cloud providers on required features and escalations.
- Participate in on-call and lead incident response.
- Develop postmortems, runbooks, and SLO processes to improve reliability and prevent recurring failures.
Requirements
- Significant software engineering experience building and operating production distributed systems.
- Proficiency in at least one systems-appropriate language, such as Go, Python, Rust, or C++.
- Deep, hands-on Kubernetes experience beyond general usage, including experience with the scheduler, controllers, apiserver, or large multi-tenant clusters.
- Ability to debug complex issues across the stack, from API behavior to node- and network-level root causes.
- Experience designing systems for reliability, correctness, and clear failure semantics.
- Strong written and verbal communication skills, including the ability to build consensus with internal stakeholders.
Preferred Qualifications
- Experience with Kubernetes internals or contributions, including kube-scheduler, the scheduling framework, apiserver, etcd, client-go, or controller-runtime.
- Experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or equivalent systems.
- Experience scaling control planes or coordination systems such as etcd, ZooKeeper, or Consul, or operating large DNS or service-mesh deployments.
- Familiarity with ML infrastructure, including GPUs, TPUs, Trainium, gang scheduling, topology-aware placement, and collective networking such as NCCL.
- Experience with GCP and/or AWS, including GKE or EKS internals and Infrastructure as Code.
- Low-level systems experience, including Linux kernel tuning, cgroups, or eBPF.
- 12 or more years of relevant industry experience, including experience leading large, ambiguous infrastructure projects.
Education And Experience
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role, as demonstrated through coursework, training, or professional experience.
- Required years of experience correlate with the internal job level requirements for the position.
Work Policy And Benefits
Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time. Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
Anthropic sponsors visas for some roles and candidates and states that it will make every reasonable effort to obtain a visa for candidates who receive an offer. The company retains an immigration lawyer to assist with this process.