Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 6
Azure @ 6
Distributed Systems @ 6
GCP @ 6
GPU
Go @ 7
IaC
InfiniBand @ 3
Kubernetes @ 6
Linux @ 6
Machine Learning @ 4
Networking @ 3
Python @ 7
Rust @ 7
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Infrastructure organization builds systems that support reliable AI research and enable Claude to scale to millions of users. The Node Infra team owns the full lifecycle of accelerator capacity, including compute ingestion and provisioning across major cloud providers and Anthropic-built datacenters, cluster scaling, and health, diagnostics, and repair automation for GPUs, TPUs, and Trainium nodes.
Responsibilities
- Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
- Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
- Design and operate systems that detect, isolate, and automatically remediate unhealthy hardware, improving fleet MTBI and minimizing stranded capacity.
- Define infrastructure architecture and ensure complex technical problems are solved directly or through collaboration with other engineers.
- Work with cloud providers and internal research, inference, and product teams on long-term compute, data, and infrastructure strategy.
- Establish and evolve operational excellence practices, including incident response, postmortems, and on-call processes.
- Support the growth of engineers through technical mentorship and coaching.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, AWS, GCP, or Azure.
- Strong proficiency in at least one systems language, such as Rust, Go, or Python.
- Proficiency with Terraform and infrastructure as code.
- Hands-on experience with machine learning accelerators, including GPUs, TPUs, or Trainium.
- Experience leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively at all levels.
- A bachelor's degree or equivalent combination of education, training, and experience in a relevant field.
Preferred Qualifications
- 12+ years of software engineering experience, including experience as a technical lead setting direction for a team.
- Experience managing hyperscale compute infrastructure with 10,000+ nodes, including capacity management and efficiency.
- Expertise in Kubernetes internals, such as the scheduler, autoscaler, kubelet, or Karpenter; cluster orchestration systems such as Mesos or Borg-like systems; or node provisioning pipelines.
- Low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
- Familiarity with high-performance networking, including EFA, RDMA, or InfiniBand, for distributed machine learning workloads.
- Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems.
- Contributions to relevant open-source projects, such as Kubernetes, the Linux kernel, or container runtimes.
- Ability to quickly understand systems design tradeoffs and track rapidly evolving software systems.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time.
More jobs at Anthropic
Developer Relations
Anthropic · San Francisco, United States, New York City, United States
USD 290,000-435,000 per year
AV Operations Specialist
Anthropic · San Francisco, United States, New York City, United States
USD 230,000-285,000 per year
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Staff Software Engineer, Claude Code
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-625,000 per year
Applied AI Architect, Enterprise Tech
Anthropic · San Francisco, United States, New York City, United States
USD 240,000-315,000 per year
Similar jobs
Senior Staff+ Software Engineer, Node Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff+ Infrastructure Engineer, Cluster Infrastructure
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Staff+ Infrastructure Engineer, Cluster Infrastructure
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year