Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 2
Data Analysis @ 3
Distributed Systems @ 3
GPU @ 2
LLM @ 3
Machine Learning
Observability @ 3
Pandas @ 3
Profiling @ 3
Python @ 5
SQL @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's inference fleet serves Claude to millions of users across its own products and the world's largest cloud platforms. The Inference System Dynamics team evaluates the system across throughput, latency, reliability, and correctness. The team measures fleet performance against theoretical performance frontiers, investigates cross-layer performance gaps, and owns correctness checks across hardware platforms and serving configurations.
Responsibilities
- Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them.
- Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead investigations when it catches a regression.
- Build observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack.
- Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and implement high-impact optimizations.
- Prioritize opportunities by impact and effort.
- Investigate latency regressions across request timing, routing, batching, server scheduling, and kernel overhead.
- Design correctness evaluation gates and regression-detection criteria for model outputs across hardware backends.
- Build performance analyses such as FLOPs funnels and models of latency, cost, batch sizing, and utilization for production autoscaling.
Requirements
- Hands-on performance engineering experience, including profiling, roofline analysis, latency and throughput optimization, and root-cause investigation in complex production systems.
- Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases.
- Solid data analysis skills, such as SQL, pandas, or similar tools, sufficient to turn raw telemetry into clear findings.
- Ability to communicate quantitative results clearly in writing and influence priorities across teams.
- Genuine interest in correctness as an engineering discipline, including numerics, evaluation design, and regression detection.
- Experience with ML systems, particularly training or inference infrastructure or LLM serving stacks, is preferred.
- Familiarity with GPU, TPU, or accelerator performance concepts, including memory bandwidth, kernel overheads, quantization, and collective communication, is preferred.
- Experience with reliability engineering for high-throughput services, including autoscaling, load balancing, request routing, and tail latency, is preferred.
- Experience with model evaluation or numerical regression-detection pipelines is preferred.
- Experience building observability or telemetry for distributed systems is preferred.
- Minimum education is a bachelor's degree or an equivalent combination of education, training, and experience. The field of study must be relevant to the role as demonstrated through coursework, training, or professional experience.
Logistics
- The role follows a location-based hybrid policy. Staff are currently expected to work from one of the company's offices at least 25% of the time, though some roles may require more office time.
- Visa sponsorship is available, although sponsorship cannot be guaranteed for every role or candidate.
- Applications are reviewed on a rolling basis; there is no application deadline.
More jobs at Anthropic
Business Systems Analyst
Anthropic · London, United Kingdom
GBP 130,000-165,000 per year
Product Manager, Business Technology
Anthropic · London, United Kingdom
GBP 190,000-240,000 per year
Incident Manager - Detection & Response
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Staff+ Software Engineer, Auth & Identity
Anthropic · New York City, United States, San Francisco, United States
USD 405,000-485,000 per year
Accounting Analytics & BI Engineer
Anthropic · San Francisco, United States
USD 220,000-270,000 per year
Similar jobs
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Technical Cyber Threat Investigator
Anthropic · Washington, United States, San Francisco, United States
USD 230,000-290,000 per year
Technical CBRN-E Threat Investigator
Anthropic · Washington, United States, San Francisco, United States
USD 230,000-290,000 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Data Management Professional - Data Engineering - Corporate Actions
Bloomberg · Princeton, United States
USD 110,000-190,000 per year