Senior Staff+ Software Engineer, Node Infra

GBP 325,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI AWS @ 6 Azure @ 6 Distributed Systems @ 6 GCP @ 6 GPU Go @ 7 IaC InfiniBand @ 3 Kubernetes @ 6 Linux @ 6 Machine Learning @ 4 Networking @ 3 Python @ 7 Rust @ 7 Terraform @ 6

Details

Anthropic's Infrastructure organization builds systems that support reliable AI research and enable Claude to scale to millions of users. The Node Infra team owns the full lifecycle of accelerator capacity, including compute ingestion and provisioning across major cloud providers and Anthropic-built datacenters, cluster scaling, and health, diagnostics, and repair automation for GPUs, TPUs, and Trainium nodes.

Responsibilities

  • Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
  • Design and operate systems that detect, isolate, and automatically remediate unhealthy hardware, improving fleet MTBI and minimizing stranded capacity.
  • Define infrastructure architecture and ensure complex technical problems are solved directly or through collaboration with other engineers.
  • Work with cloud providers and internal research, inference, and product teams on long-term compute, data, and infrastructure strategy.
  • Establish and evolve operational excellence practices, including incident response, postmortems, and on-call processes.
  • Support the growth of engineers through technical mentorship and coaching.

Requirements

  • Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, AWS, GCP, or Azure.
  • Strong proficiency in at least one systems language, such as Rust, Go, or Python.
  • Proficiency with Terraform and infrastructure as code.
  • Hands-on experience with machine learning accelerators, including GPUs, TPUs, or Trainium.
  • Experience leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels.
  • A bachelor's degree or equivalent combination of education, training, and experience in a relevant field.

Preferred Qualifications

  • 12+ years of software engineering experience, including experience as a technical lead setting direction for a team.
  • Experience managing hyperscale compute infrastructure with 10,000+ nodes, including capacity management and efficiency.
  • Expertise in Kubernetes internals, such as the scheduler, autoscaler, kubelet, or Karpenter; cluster orchestration systems such as Mesos or Borg-like systems; or node provisioning pipelines.
  • Low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
  • Familiarity with high-performance networking, including EFA, RDMA, or InfiniBand, for distributed machine learning workloads.
  • Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems.
  • Contributions to relevant open-source projects, such as Kubernetes, the Linux kernel, or container runtimes.
  • Ability to quickly understand systems design tradeoffs and track rapidly evolving software systems.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time.

More jobs at Anthropic

Similar jobs