Engineering Manager, Inference Infrastructure

USD 405,000-625,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

API Engineering Management @ 4 Hiring @ 4 Kubernetes @ 4 LLM @ 4 Machine Learning @ 6 Networking @ 7

Details

Anthropic is seeking an engineering manager to lead the team responsible for the control plane coordinating its inference fleet and the inference request path. The team determines where requests are served and how much capacity each model requires, with a focus on throughput, reliability, latency, utilization, and cost.

This is a deeply technical role leading ML platform, infrastructure, and distributed-systems engineers. The manager will work with teams responsible for ML internals, inference engines, performance, capacity, product, and cloud infrastructure.

Responsibilities

  • Own the technical roadmap for coordinating the inference fleet, including traffic routing, capacity placement, cache placement, demand responsiveness, and synchronization protocols between the control plane and inference engines.
  • Partner with product, inference engine, performance, and capacity teams to identify and ship measurable improvements in throughput, latency, utilization, and cost.
  • Establish quantitative modeling practices and measure expected and actual system impact.
  • Set technical strategy across heterogeneous hardware, multiple cloud providers, and all serving surfaces.
  • Lead operational practices including on-call rotations, incident response, postmortem reviews, and deployment safety.
  • Create clarity across the API surface, inference engines, capacity planning, and cloud deployment teams.
  • Develop and retain existing teams, hire engineers to a high technical bar, and coach engineers through shifting priorities.
  • Shape team structure as the organization grows and develop leads who can own problem areas.
  • Unblock critical initiatives and synthesize design debates when necessary.

Requirements

  • Engineering management experience leading teams responsible for critical-path production infrastructure at scale.
  • Deep systems experience in areas such as load balancing, scheduling, cluster orchestration, autoscaling, cache-coherent distributed state, high-performance networking, or similar systems.
  • Ability to make architectural decisions about large-scale fleet coordination and evaluate engineers working at the kernel and framework levels.
  • Experience delivering performance or efficiency improvements in large-scale systems and quantifying their latency and cost impact.
  • Experience operating production infrastructure with on-call, incident response, capacity events, and deployment discipline.
  • Results-oriented and impact-driven approach, with the ability to balance throughput, latency, cost, stability, launch timelines, and feature velocity.
  • Ability to build strong relationships across team boundaries.
  • Curiosity about machine learning systems and transformer inference.

Preferred Qualifications

  • 5+ years of engineering management experience.
  • Experience with LLM inference serving, including KV caching, continuous batching, request scheduling, and prefill/decode disaggregation.
  • Experience with cluster schedulers, autoscalers, load balancers, service meshes, or fleet control planes at scale, including Kubernetes internals, Borg-style systems, or equivalents.
  • Experience operating workloads across multiple clouds or partner platforms.
  • Familiarity with heterogeneous accelerator fleets and the effect of hardware differences on workload placement and rollout sequencing.
  • Experience leading teams at supercomputing or hyperscaler infrastructure scale.
  • Experience leading multiple teams or managing rapid-growth periods involving hiring, onboarding, and team splits.

Education and Logistics

  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
  • Required field of study: A field relevant to the role, demonstrated through coursework, training, or professional experience.
  • The role follows a location-based hybrid policy, with staff expected to work from an Anthropic office at least 25% of the time; some roles may require more office time.
  • Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for candidates receiving an offer, although sponsorship cannot be guaranteed for every role or candidate.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs