Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 6
Azure @ 6
Distributed Systems @ 6
GCP @ 6
GPU
Go @ 7
IaC
InfiniBand @ 3
Kubernetes @ 6
Linux @ 6
Machine Learning @ 4
Networking @ 3
Python @ 7
Rust @ 7
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Infrastructure organization builds systems that support reliable AI research and enable Claude to scale to millions of users. The Node Infra team owns the full lifecycle of accelerator capacity, including compute ingestion and provisioning across major cloud providers and Anthropic-built datacenters, cluster scaling, and health, diagnostics, and repair automation for GPUs, TPUs, and Trainium nodes.
Responsibilities
- Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
- Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
- Design and operate systems that detect, isolate, and automatically remediate unhealthy hardware, improving fleet MTBI and minimizing stranded capacity.
- Define infrastructure architecture and ensure complex technical problems are solved directly or through collaboration with other engineers.
- Work with cloud providers and internal research, inference, and product teams on long-term compute, data, and infrastructure strategy.
- Establish and evolve operational excellence practices, including incident response, postmortems, and on-call processes.
- Support the growth of engineers through technical mentorship and coaching.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, AWS, GCP, or Azure.
- Strong proficiency in at least one systems language, such as Rust, Go, or Python.
- Proficiency with Terraform and infrastructure as code.
- Hands-on experience with machine learning accelerators, including GPUs, TPUs, or Trainium.
- Experience leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively at all levels.
- A bachelor's degree or equivalent combination of education, training, and experience in a relevant field.
Preferred Qualifications
- 12+ years of software engineering experience, including experience as a technical lead setting direction for a team.
- Experience managing hyperscale compute infrastructure with 10,000+ nodes, including capacity management and efficiency.
- Expertise in Kubernetes internals, such as the scheduler, autoscaler, kubelet, or Karpenter; cluster orchestration systems such as Mesos or Borg-like systems; or node provisioning pipelines.
- Low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
- Familiarity with high-performance networking, including EFA, RDMA, or InfiniBand, for distributed machine learning workloads.
- Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems.
- Contributions to relevant open-source projects, such as Kubernetes, the Linux kernel, or container runtimes.
- Ability to quickly understand systems design tradeoffs and track rapidly evolving software systems.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time.
More jobs at Anthropic
Product Manager, Growth
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 385,000-460,000 per year
Staff+ Research Engineer, RL Data Platform
Anthropic · New York City, United States, San Francisco, United States
USD 500,000-850,000 per year
Engineering Manager, Hardware Platform Security
Anthropic · San Francisco, United States, Seattle, United States
USD 485,000-625,000 per year
Staff+ Software Engineer, RL Data Platform
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-405,000 per year
Applied AI Architect, Strategic Enterprise Tech
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 240,000-315,000 per year
Similar jobs
Senior Staff+ Software Engineer, Node Infra
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Software Engineer, Infrastructure (Distributed Systems)
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Staff+ Infrastructure Engineer, Cluster Infrastructure
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff Platform Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year