Staff+ Software Engineer, Inference Runtime

USD 405,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 4 AWS @ 6 CI/CD @ 4 CUDA @ 6 Communication @ 7 Debugging @ 6 Distributed Systems @ 7 GPU @ 6 Kubernetes @ 3 Machine Learning Profiling @ 6 Python Rust

Details

Anthropic's Inference organization serves Claude to millions of users and enterprise customers with the speed, reliability, and efficiency required for frontier AI. The Inference Runtime team owns the shared, accelerator-agnostic core of the inference serving stack across GPUs, TPUs, and Trainium.

This is a senior individual contributor role with broad technical ownership. The Staff Engineer will set technical direction for the runtime architecture, release and validation systems, and engineering workflows. The role partners across Inference on boundaries, prioritization, and tradeoffs across heterogeneous accelerator platforms, while working alongside the Engineering Manager, who owns hiring and people development.

Responsibilities

  • Set technical direction for the team and own the architecture and roadmap for the shared runtime of the inference serving stack.
  • Own and evolve the accelerator-agnostic runtime, including its interfaces, internal boundaries, and build structure.
  • Perform hands-on work in a performance-sensitive Rust and Python codebase.
  • Keep platform expansion costs low by ensuring new models and deployment targets pay only for their own specialization and that edge cases integrate cleanly with the core.
  • Drive efficient accelerator usage, including utilization, scheduling, and memory management across GPU, TPU, and Trainium.
  • Build the runtime's validation surface around partitioned builds, change-scoped testing, and canary, shadow, and rollback mechanisms.
  • Act as a technical counterpart to Anthropic's central Infrastructure organization on compilers, build systems, and toolchains.
  • Contribute Inference's performance and correctness requirements and make build-versus-adopt decisions.
  • Mentor engineers through design reviews, code reviews, and direct collaboration.
  • Represent the team in cross-organizational efforts spanning serving, scaling, and accelerator teams.

Requirements

  • Deep background in systems engineering or ML infrastructure.
  • Ability to work hands-on with performance profiling, latency and throughput optimization, and systems debugging at scale.
  • Significant depth in at least one accelerator ecosystem: CUDA/GPU, TPU, or Trainium/AWS Neuron.
  • Significant software engineering experience and a strong background in high-performance, large-scale distributed systems serving millions of users.
  • Track record of defining and using engineering metrics to drive improvement, including platform SLOs, escape rates, release times, latency, or throughput.
  • Experience driving technical alignment across organizational boundaries and influencing technical direction without formal authority.
  • Strong written and verbal communication skills.
  • A bachelor's degree or equivalent combination of education, training, and/or experience in a field relevant to the role, as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • 8+ years of software engineering experience, including significant time as the technical lead or anchor on a platform, inference runtime, or ML infrastructure team.
  • Experience with ML compiler toolchains such as XLA, Triton, or NeuronX, or with accelerator driver and firmware management at scale.
  • Experience operating production as a validation surface at scale, including shadow traffic, canary populations, automated baseline comparison, and fast rollback.
  • Experience with deterministic or simulation-based testing for hardware-dependent systems.
  • Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware.
  • Familiarity with Kubernetes-based development and job scheduling environments.
  • Prior tech lead experience on a developer productivity or platform engineering team at a fast-growing AI/ML company.

Compensation

  • Annual salary: $405,000–$485,000 USD.

Logistics

  • Remote-friendly, with travel required.
  • Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
  • Anthropic sponsors visas and makes reasonable efforts to obtain a visa when an offer is made, although sponsorship is not successful for every role and candidate.
  • Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs