Staff + Senior Software Engineer, Cloud Inference

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

API @ 6 AWS @ 4 Azure @ 4 CI/CD Distributed Systems @ 7 GCP @ 4 IaC Kubernetes @ 6 LLM Machine Learning @ 4 Networking @ 4 Observability Python @ 6 Rust @ 6 Security @ 6

Details

The Cloud Inference team scales and optimizes Claude across AWS, GCP, Azure, and future cloud service providers. The team owns the end-to-end Claude product on each cloud platform, including API integration, intelligent request routing, inference execution, capacity management, and operations.

The role focuses on building reliable, cost-effective, large-scale backend services and infrastructure for Claude across cloud providers, supporting the launch of new frontier models and features while meeting rigorous safety, performance, and security standards.

Responsibilities

  • Design, build, and own backend services and infrastructure serving Claude across multiple cloud service providers, accounting for differences in compute hardware, networking, APIs, and operational models.
  • Collaborate with internal inference, product API, systems, and security teams, as well as cloud service provider partners, to build serving stacks, resolve operational issues, and influence provider roadmaps.
  • Build and evolve CI/CD automation, including validation and deployment pipelines for reliably shipping new model versions to millions of users.
  • Design interfaces and tooling abstractions across cloud providers to enable cost-effective inference management, scale across providers, and reduce platform-specific complexity.
  • Contribute to capacity planning, autoscaling, and workload-routing strategies that match supply with demand and direct requests to cost-effective accelerators and regions.
  • Analyze observability data to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation using production workloads.

Requirements

  • Significant software engineering experience and a strong background in high-performance, large-scale distributed systems serving millions of users.
  • Experience building or operating services on at least one major cloud platform: AWS, GCP, or Azure.
  • Exposure to Kubernetes, infrastructure as code, or container orchestration.
  • Curiosity about large language model serving; prior inference or machine learning experience is not required.
  • Experience collaborating cross-functionally with internal teams and external partners.
  • Ability to quickly learn new technologies, hardware platforms, and provider ecosystems.
  • High autonomy and end-to-end ownership.

Preferred Qualifications

  • Experience working with cloud service providers to scale infrastructure or products across multiple platforms, including differences in networking, security, privacy, billing, and managed services.
  • Hands-on experience with capacity management, cost optimization, or resource planning at scale across heterogeneous environments.
  • Understanding of multi-region deployments, geographic routing, and global traffic management.
  • Proficiency in Python or Rust.

Education and Experience

  • A bachelor's degree or equivalent combination of education, training, and experience is required.
  • The required field of study should be relevant to the role through coursework, training, or professional experience.
  • Required years of experience correlate with the internal job level.

Benefits and Logistics

  • Annual salary: $320,000–$485,000 USD.
  • Hybrid policy: Staff are expected to work from one of the offices at least 25% of the time, although some roles may require more office time.
  • Anthropic sponsors visas where possible and makes reasonable efforts to obtain visas for candidates who receive an offer.
  • Benefits include competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.

More jobs at Anthropic

Similar jobs