Staff+ Software Engineer, Capacity Engineering

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

API @ 4 Azure @ 7 Data Engineering Data Pipelines DevOps @ 7 GPU @ 3 Grafana @ 4 Kubernetes Machine Learning @ 4 Observability @ 4 Planning @ 4 Profiling @ 4 Prometheus @ 4 Python @ 6 SQL @ 6

Details

Anthropic’s Capacity Engineering team builds the data, tooling, and operational systems used to account for, utilize, allocate, and plan infrastructure resources across accelerator families, CPU families, and cloud providers. The role focuses on production systems spanning data engineering, systems engineering, observability, capacity planning, efficiency measurement, attribution, and forecasting.

This is a pipeline role feeding four areas: data platform, planning, efficiency, and attribution and forecasting. Depending on background and business priorities, the engineer will focus primarily in one area while collaborating across all four.

Responsibilities

  • Build the planning and allocation stack used by leadership, engineering teams, and schedulers, including cross-region and cross-provider placement, guardrails, queueing, and occupancy KPIs.
  • Drive efficiency programs involving stranding and rightsizing, unused-capacity recovery, and job-level utilization across training, inference, and evaluation workloads.
  • Establish per-configuration baselines and work with system-owning teams to improve utilization.
  • Own attribution and forecasting by reconciling billing across more than ten providers with telemetry and internal systems, attributing spend to workloads, and converting demand signals and research roadmaps into compute plans and supply pipelines.
  • Build data pipelines that ingest occupancy, utilization, and cost data from a diversifying infrastructure fleet into BigQuery, with ownership of completeness, latency SLOs, and gap detection.
  • Integrate data from new cloud and infrastructure providers, including billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry.
  • Operate Kubernetes-native systems at scale, including collection agents, workload labeling, taints, reservations, and scheduling behavior.
  • Build observability tooling and performance instrumentation for fleet health and workload efficiency.
  • Gather requirements, define schema contracts, design self-service data products, and support consumers ranging from research engineers to finance and company leadership.
  • Operate load-bearing systems with on-call responsibilities and SLOs.

Requirements

  • Strong track record building and operating production systems in a hands-on engineering or DevOps-oriented role.
  • Production-quality Python and SQL skills. Pipeline code is primarily Python, while the presentation layer uses BigQuery SQL, including table-valued functions and views.
  • Deep experience with at least one major cloud provider: Amazon Web Services, Google Cloud, or Microsoft Azure.
  • Experience with observability tooling, including Prometheus, PromQL, and Grafana; recording rules and monitoring relied upon by engineering teams.
  • Ability to gather requirements independently and work across organizational boundaries in an ambiguous environment with limited direction.
  • Bachelor’s degree or equivalent combination of education, training, and experience in a field relevant to the role, as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Capacity planning, resource management, or cost attribution experience at a hyperscaler or in a large-scale machine learning environment.
  • Scheduling and packing efficiency experience or profiling-driven optimization of large distributed workloads.
  • Multi-cloud data ingestion experience, especially normalizing billing exports and telemetry from providers with different billing arrangements.
  • Total cost of ownership and forecasting experience, including analysis of infrastructure growth against business drivers.
  • Accelerator infrastructure familiarity, including GPU metrics such as DCGM, TPU utilization, Trainium power and utilization metrics, or machine learning training and inference systems at the hardware level.
  • Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability.
  • Storage efficiency, retention, and lifecycle program experience at exabyte scale.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Staff are currently expected to work from an Anthropic office at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for candidates who receive an offer, with assistance from an immigration lawyer.

More jobs at Anthropic

Similar jobs