Performance Engineer, Inference Systems

USD 350,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Used Tools & Technologies

Machine Learning

Required Skills & Competences

Python @ 5 SQL @ 3 Distributed Systems @ 3 Communication @ 2 Data Analysis @ 3 LLM @ 1 Pandas @ 3 GPU @ 2 Observability @ 3 AI @ 3 Profiling @ 3

Details

Anthropic’s inference fleet serves Claude to millions of users across our own products and the world's largest cloud platforms. The Inference System Dynamics team is responsible for understanding the whole serving stack (accelerator kernels, model servers, distributed routing, autoscaling, capacity management) and holding it to a high bar across throughput, latency, reliability, and correctness. The team instruments and models components, runs cross-layer investigations, and partners with owning teams to land optimizations.

Responsibilities

  • Run cross-layer performance investigations across throughput, latency, and reliability; size the gap between actual fleet performance and theoretical rooflines; identify root causes and quantify the value of closing them
  • Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations; lead investigations when it catches a regression
  • Build observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack
  • Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land high-impact optimizations
  • Ruthlessly stack-rank opportunities by impact and effort and decline low-leverage work

Requirements

Minimum Qualifications:

  • Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems
  • Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write
  • Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings
  • Ability to communicate quantitative results clearly in writing to influence priorities across teams
  • Genuine interest in correctness as an engineering discipline: numerics, evaluation design, regression detection

Preferred Qualifications:

  • Experience with ML systems, especially training or inference infrastructure or general LLM serving stacks; direct large-scale inference experience is a strong plus
  • Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication)
  • Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency
  • Experience with model evaluation or numerical regression-detection pipelines
  • Experience building observability or telemetry for distributed systems
  • Comfortable having impact through influence and evidence rather than direct ownership

Representative Projects

  • Trace a 350ms latency gap on a new accelerator platform from end-to-end request timing down to a server scheduling overhead, quantify the win, and land the fix
  • Redesign the correctness eval gate to reliably catch real model-output regressions across hardware backends
  • Build a FLOPs funnel that breaks down where compute goes across the fleet, exposing gaps between achieved throughput and kernel rooflines
  • Root-cause numerical divergence between hardware platforms to a kernel change and define acceptance thresholds
  • Model latency–cost impacts of changing batch-sizing and utilization targets and convert the result into autoscaler signals

Compensation

  • Annual Salary: $350,000 - $850,000 USD

Logistics

  • Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Will correlate with internal job level requirements
  • Location-based hybrid policy: staff expected to be in one of the offices at least 25% of the time (some roles may require more time in office)
  • Visa sponsorship: Anthropic states they sponsor visas and will make reasonable efforts to support candidates offered roles (they retain an immigration lawyer)

Benefits

  • Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration

How we're different

  • Work as a cohesive team on a few large-scale research efforts with a focus on impact. Frequent research discussions and emphasis on communication skills. The team values empirical, large-scale AI research and diverse perspectives.

Application notes

  • Applications reviewed on a rolling basis. Guidance on candidate AI usage and application instructions are provided on Anthropic’s careers pages.