Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 6
CI/CD @ 4
CUDA @ 6
Communication @ 7
Debugging @ 6
Distributed Systems @ 7
GPU @ 6
Kubernetes @ 3
Machine Learning
Profiling @ 6
Python
Rust
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Inference organization serves Claude to millions of users and enterprise customers with the speed, reliability, and efficiency required for frontier AI. The Inference Runtime team owns the shared, accelerator-agnostic core of the inference serving stack across GPUs, TPUs, and Trainium.
This is a senior individual contributor role with broad technical ownership. The Staff Engineer will set technical direction for the runtime architecture, release and validation systems, and engineering workflows. The role partners across Inference on boundaries, prioritization, and tradeoffs across heterogeneous accelerator platforms, while working alongside the Engineering Manager, who owns hiring and people development.
Responsibilities
- Set technical direction for the team and own the architecture and roadmap for the shared runtime of the inference serving stack.
- Own and evolve the accelerator-agnostic runtime, including its interfaces, internal boundaries, and build structure.
- Perform hands-on work in a performance-sensitive Rust and Python codebase.
- Keep platform expansion costs low by ensuring new models and deployment targets pay only for their own specialization and that edge cases integrate cleanly with the core.
- Drive efficient accelerator usage, including utilization, scheduling, and memory management across GPU, TPU, and Trainium.
- Build the runtime's validation surface around partitioned builds, change-scoped testing, and canary, shadow, and rollback mechanisms.
- Act as a technical counterpart to Anthropic's central Infrastructure organization on compilers, build systems, and toolchains.
- Contribute Inference's performance and correctness requirements and make build-versus-adopt decisions.
- Mentor engineers through design reviews, code reviews, and direct collaboration.
- Represent the team in cross-organizational efforts spanning serving, scaling, and accelerator teams.
Requirements
- Deep background in systems engineering or ML infrastructure.
- Ability to work hands-on with performance profiling, latency and throughput optimization, and systems debugging at scale.
- Significant depth in at least one accelerator ecosystem: CUDA/GPU, TPU, or Trainium/AWS Neuron.
- Significant software engineering experience and a strong background in high-performance, large-scale distributed systems serving millions of users.
- Track record of defining and using engineering metrics to drive improvement, including platform SLOs, escape rates, release times, latency, or throughput.
- Experience driving technical alignment across organizational boundaries and influencing technical direction without formal authority.
- Strong written and verbal communication skills.
- A bachelor's degree or equivalent combination of education, training, and/or experience in a field relevant to the role, as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- 8+ years of software engineering experience, including significant time as the technical lead or anchor on a platform, inference runtime, or ML infrastructure team.
- Experience with ML compiler toolchains such as XLA, Triton, or NeuronX, or with accelerator driver and firmware management at scale.
- Experience operating production as a validation surface at scale, including shadow traffic, canary populations, automated baseline comparison, and fast rollback.
- Experience with deterministic or simulation-based testing for hardware-dependent systems.
- Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware.
- Familiarity with Kubernetes-based development and job scheduling environments.
- Prior tech lead experience on a developer productivity or platform engineering team at a fast-growing AI/ML company.
Compensation
- Annual salary: $405,000–$485,000 USD.
Logistics
- Remote-friendly, with travel required.
- Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and makes reasonable efforts to obtain a visa when an offer is made, although sponsorship is not successful for every role and candidate.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
More jobs at Anthropic
Business Systems Analyst
Anthropic · London, United Kingdom
GBP 130,000-165,000 per year
Product Manager, Business Technology
Anthropic · London, United Kingdom
GBP 190,000-240,000 per year
Incident Manager - Detection & Response
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Staff+ Software Engineer, Auth & Identity
Anthropic · New York City, United States, San Francisco, United States
USD 405,000-485,000 per year
Accounting Analytics & BI Engineer
Anthropic · San Francisco, United States
USD 220,000-270,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year