Performance Engineer, Inference Systems

USD 350,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Communication @ 2 Data Analysis @ 3 Distributed Systems @ 3 GPU @ 2 LLM @ 3 Machine Learning Observability @ 3 Pandas @ 3 Profiling @ 3 Python @ 5 SQL @ 3

Details

Anthropic's inference fleet serves Claude to millions of users across its own products and the world's largest cloud platforms. The Inference System Dynamics team evaluates the system across throughput, latency, reliability, and correctness. The team measures fleet performance against theoretical performance frontiers, investigates cross-layer performance gaps, and owns correctness checks across hardware platforms and serving configurations.

Responsibilities

  • Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them.
  • Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead investigations when it catches a regression.
  • Build observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack.
  • Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and implement high-impact optimizations.
  • Prioritize opportunities by impact and effort.
  • Investigate latency regressions across request timing, routing, batching, server scheduling, and kernel overhead.
  • Design correctness evaluation gates and regression-detection criteria for model outputs across hardware backends.
  • Build performance analyses such as FLOPs funnels and models of latency, cost, batch sizing, and utilization for production autoscaling.

Requirements

  • Hands-on performance engineering experience, including profiling, roofline analysis, latency and throughput optimization, and root-cause investigation in complex production systems.
  • Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases.
  • Solid data analysis skills, such as SQL, pandas, or similar tools, sufficient to turn raw telemetry into clear findings.
  • Ability to communicate quantitative results clearly in writing and influence priorities across teams.
  • Genuine interest in correctness as an engineering discipline, including numerics, evaluation design, and regression detection.
  • Experience with ML systems, particularly training or inference infrastructure or LLM serving stacks, is preferred.
  • Familiarity with GPU, TPU, or accelerator performance concepts, including memory bandwidth, kernel overheads, quantization, and collective communication, is preferred.
  • Experience with reliability engineering for high-throughput services, including autoscaling, load balancing, request routing, and tail latency, is preferred.
  • Experience with model evaluation or numerical regression-detection pipelines is preferred.
  • Experience building observability or telemetry for distributed systems is preferred.
  • Minimum education is a bachelor's degree or an equivalent combination of education, training, and experience. The field of study must be relevant to the role as demonstrated through coursework, training, or professional experience.

Logistics

  • The role follows a location-based hybrid policy. Staff are currently expected to work from one of the company's offices at least 25% of the time, though some roles may require more office time.
  • Visa sponsorship is available, although sponsorship cannot be guaranteed for every role or candidate.
  • Applications are reviewed on a rolling basis; there is no application deadline.

More jobs at Anthropic

Similar jobs