Senior Staff+ Software Engineer, Node Infra

USD 405,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI AWS @ 6 Azure @ 6 Distributed Systems @ 6 GCP @ 6 GPU Go @ 7 IaC InfiniBand @ 3 Kubernetes @ 6 Linux @ 6 Machine Learning @ 4 Networking @ 3 Python @ 7 Rust @ 7 Terraform @ 6

Details

Anthropic's Infrastructure organization builds systems that enable reliable AI research and the operation of Claude at scale. The Node Infra team owns the full lifecycle of accelerator capacity, including compute ingestion and provisioning across major cloud providers and custom-built datacenters, cluster scaling, and health, diagnostics, and repair automation for GPUs, TPUs, and Trainium nodes.

Responsibilities

  • Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
  • Design and operate systems that automatically detect, isolate, and remediate unhealthy hardware, improving fleet MTBI and minimizing stranded capacity.
  • Define infrastructure architecture and ensure difficult technical problems are solved directly or through effective collaboration.
  • Work with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy.
  • Establish and evolve operational excellence practices, including incident response, postmortem culture, and on-call processes.
  • Support the growth of engineers through technical mentorship and coaching.

Requirements

  • Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, and AWS, GCP, or Azure.
  • Strong proficiency in at least one systems language, such as Rust, Go, or Python.
  • Proficiency with Terraform and infrastructure as code.
  • Hands-on experience with machine learning accelerators, including GPUs, TPUs, or Trainium.
  • Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels.
  • Bachelor’s degree or an equivalent combination of education, training, and experience in a relevant field.

Preferred Qualifications

  • 12+ years of software engineering experience, including experience as a technical lead setting direction for a team.
  • Experience managing large-scale compute infrastructure at hyperscale, including 10,000+ nodes, capacity management, and efficiency.
  • Depth in Kubernetes internals, including the scheduler, autoscaler, kubelet, or Karpenter; cluster orchestration systems such as Mesos or Borg-like systems; or node provisioning pipelines.
  • Low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
  • Familiarity with high-performance networking, including EFA, RDMA, or InfiniBand, for distributed machine learning workloads.
  • Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems.
  • Contributions to relevant open-source projects such as Kubernetes, the Linux kernel, or container runtimes.
  • Ability to quickly understand systems design tradeoffs and track rapidly evolving software systems.

Compensation

Annual salary: $405,000–$485,000 USD.

Benefits and Logistics

  • Location-based hybrid policy requiring staff to work from one of the offices at least 25% of the time; some roles may require more office time.
  • Anthropic sponsors visas and makes reasonable efforts to obtain a visa for successful candidates, with support from an immigration lawyer.
  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Collaborative office space.

More jobs at Anthropic

Similar jobs