AI Infrastructure Operations, Demand Planning

USD 320,000-405,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

HPC @ 3 Observability @ 3 Python @ 3 Reporting @ 3 SQL @ 3

Details

Anthropic’s Capacity Engineering organization owns the data, tooling, and systems used to plan, measure, and maximize utilization across infrastructure fleets spanning accelerator and CPU families, clouds, neoclouds, and on-premises sites. This role is part of the Planning pillar on the Demand Planning team and works with research engineering, pretraining, inference, compute supply, finance, and external vendors.

The role converts demand forecasts into per-tranche infrastructure requirements and owns the integrated schedule and system of record for each tranche, from contract through reservation, ingestion, cluster health, and occupancy. It also drives delivery improvements through contract changes, automation, and forecast feedback.

Responsibilities

  • Convert Demand Planning forecasts and input from research, pretraining, and inference planners into accelerator, interconnect, region, supporting-resource, and date requirements for each tranche.
  • Represent tranche requirements in sourcing negotiations and data center build reviews, including contractual terms that affect delivery dates.
  • Qualify tranches for deliverability before signature, ensuring they can be schedulable, healthy, and instrumented in the target region and timeframe, with storage, egress, and identity available.
  • Track forecast versus delivered capacity across shape, region, and timing; publish variance and feed it back into Demand Planning and future contracts.
  • Define the canonical contract-to-occupied state machine with explicit entry and exit criteria for each stage, and make it a first-class object in the capacity data layer.
  • Run multiple bring-ups in parallel across new cloud regions, on-premises sites, and neocloud blocks.
  • Maintain an integrated schedule covering provider milestones, cluster creation, network turn-up, storage readiness, health burn-in, and first-workload landing.
  • Drive readiness automation and ensure capacity systems are integrated and scaled for all new capacity from contract through ingestion.
  • Instrument and publish time-to-occupied and paid-idle dollars per tranche, with executive-level reporting on portfolio status, tradeoffs, and risk.

Requirements

  • Significant experience delivering large-scale infrastructure, including cloud regions, accelerator clusters, HPC systems, or bare-metal fleets, at multi-region scale or at least 10,000 accelerators or equivalent CPU/storage capacity.
  • Technical range spanning cluster orchestration and node health through telemetry and planning tables, with the ability to debug discrepancies between them.
  • SQL skills and enough Python to answer questions independently and build reporting.
  • Bachelor’s degree in a technical field or an equivalent engineering track record.

Preferred Qualifications

  • Experience with reserved-capacity onboarding, private offers, or capacity commitments with cloud or neocloud providers.
  • Demand-planning experience, including challenging forecasts, translating them into per-tranche requirements, and feeding delivery variance back into forecasts.
  • Data center or colocation delivery experience, including power and space planning, network turn-up, site acceptance, and vendor management.
  • Experience with accelerator health and burn-in, collective-communications sanity testing, or fleet-health SLOs.
  • Experience building systems of record or lifecycle services for infrastructure assets.
  • Experience onboarding a new hardware generation into an existing scheduler and observability stack.

Education and Logistics

  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
  • Minimum years of experience: Experience requirements correlate with the internal job level.
  • Location-based hybrid policy: Staff are currently expected to work from one of Anthropic’s offices at least 25% of the time, though some roles may require more office time.
  • Anthropic sponsors visas when possible and makes reasonable efforts to obtain a visa for candidates who receive an offer, with assistance from an immigration lawyer.

Compensation

  • Annual salary: $320,000–$405,000 USD.

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs