Staff + Senior Software Engineer, Cloud Inference Launch Engineering

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

API AWS @ 4 Azure @ 4 CI/CD Distributed Systems @ 7 GCP @ 4 IaC Kubernetes @ 6 LLM @ 7 Machine Learning @ 7 Networking @ 4 Observability Python @ 6 Rust @ 6 Security @ 4

Details

Anthropic's Cloud Inference team scales and optimizes Claude for developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers. The team owns the end-to-end product of Claude on each cloud platform, including API integration, intelligent request routing, inference execution, capacity management, and day-to-day operations.

The model and inference launch team owns the validation pipeline for the inference server and load balancer across cloud platforms. The team ensures that model launches, performance improvements, and safeguard integrations reach cloud platforms with correctness, performance, and reliability intact.

Responsibilities

  • Support frontier model launches by bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with Anthropic's first-party platform.
  • Work with the core inference team to bring new inference features, such as structured sampling and prompt caching, to cloud platforms and own the platform-specific integrations required for production.
  • Identify and resolve differences between first-party and cloud service provider inference, including configuration drift, observability gaps, deployment patterns, and difficult cross-platform bugs.
  • Design, build, and own CI/CD infrastructure for the inference server and load balancer across cloud platforms.
  • Implement shadow traffic, performance baselines for throughput and latency, and correctness checks to catch regressions before production.
  • Reduce merge-to-production cycle time by making validation faster, more parallel, and cost-effective enough to run on constrained accelerator pools without compromising reliability.
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation using real-world production workloads.

Requirements

  • Strong interest in LLM serving; prior inference or machine learning experience is not required.
  • Significant software engineering experience and a strong background in high-performance, large-scale distributed systems serving millions of users.
  • Track record of building automation or test infrastructure that measurably improved release velocity or reliability.
  • Experience building or operating services on at least one major cloud platform: AWS, GCP, or Azure.
  • Exposure to Kubernetes, infrastructure as code, or container orchestration.
  • Ability to collaborate cross-functionally with internal teams and external partners.
  • Ability to quickly learn new technologies, hardware platforms, and provider ecosystems.
  • High autonomy and end-to-end ownership, including work outside the formal job description.

Preferred Qualifications

  • LLM inference optimization, batching, and caching strategies.
  • Capacity-constrained scheduling or shared-resource test infrastructure.
  • Understanding of multi-region deployments, request routing, load balancing, and global traffic management.
  • Experience working with cloud service provider partner teams to scale infrastructure across multiple platforms, including differences in networking, security, privacy, and managed services.
  • Proficiency in Python or Rust.

Education and Experience

  • Bachelor's degree or an equivalent combination of education, training, and experience.
  • Relevant field of study demonstrated through coursework, training, or professional experience.
  • Required years of experience correlate with the internal job level requirements.

Benefits

  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Collaborative office space.
  • Anthropic sponsors visas and makes reasonable efforts to obtain a visa for candidates when an offer is made, although sponsorship is not available for every role and candidate.

The role follows a location-based hybrid policy. Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.

More jobs at Anthropic

Similar jobs