Staff + Senior Software Engineer, Inference

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI AWS @ 4 Algorithms Azure @ 4 Distributed Systems @ 4 Kubernetes @ 4 LLM Machine Learning @ 4 Observability Python @ 6 Rust @ 6

Details

Anthropic’s Inference team builds and maintains the critical systems that serve Claude to millions of users worldwide. The team operates compute-agnostic inference deployments and owns the stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.

The role focuses on maximizing compute efficiency while enabling research through high-performance inference infrastructure. The work involves complex distributed systems challenges across multiple accelerator families, emerging AI hardware, and multiple cloud platforms.

Responsibilities

  • Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
  • Develop resilient and flexible systems that adapt in real time to real-world events.
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
  • Maximize compute efficiency across the fleet through autoscaling and orchestration of production, research, and experimental workloads.
  • Build and operate production-grade deployment pipelines for releasing new models to users.
  • Provide high-performance inference infrastructure for researchers developing next-generation models.
  • Integrate new AI accelerator platforms and support inference for new model architectures.
  • Design intelligent routing algorithms that optimize request distribution across accelerators in different environments.
  • Autoscale the compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
  • Build reliable deployment pipelines for releasing new models to millions of users.
  • Contribute to new inference features and support inference for new model architectures.
  • Analyze observability data to tune performance based on real-world production workloads.
  • Manage multi-region deployments and geographic routing for global customers.

Requirements

Minimum Qualifications

  • Significant software engineering experience, particularly with distributed systems.
  • Results-oriented approach with a bias toward flexibility and impact.
  • Willingness to take on work outside the formal job description when needed.
  • Desire to learn more about machine learning systems and infrastructure.
  • Ability to thrive in environments where technical excellence drives business results and research breakthroughs.
  • Interest in the societal impacts of the work.
  • Bachelor’s degree or an equivalent combination of education, training, and experience.
  • Education or professional experience in a field relevant to the role.

Preferred Qualifications

  • Experience with high-performance, large-scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience with load balancing, request routing, or traffic management systems.
  • Familiarity with large language model inference optimization, batching, and caching strategies.
  • Experience with Kubernetes and cloud infrastructure, including AWS, Google Cloud Platform, or Azure.
  • Proficiency in Python or Rust.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration. Anthropic also states that it sponsors visas and retains an immigration lawyer to assist with immigration, although sponsorship is not guaranteed for every role or candidate.

More jobs at Anthropic

Similar jobs