Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
GPU
HPC @ 3
Observability @ 3
Pandas @ 3
Python @ 3
Reinforcement Learning
SQL @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's infrastructure fleet spans a growing set of clouds, neoclouds, and on-premises sites. The Capacity Engineering team connects capacity planning, supply management, capacity delivery, and utilization. The team partners on supply deals, wires telemetry, owns the canonical capacity data layer, and builds planning and enforcement tools for research and product teams.
As an Infrastructure Capacity Planner on the Demand Planning team, you will own the medium-range demand forecast for resource classes including accelerators, CPU, storage, network, and managed services. The forecast will support sourcing and purchasing, prioritization, efficiency targets, and workload optimization.
Responsibilities
- Build and own the medium-range, multi-resource demand forecast for:
- Accelerators by chip and interconnect class
- CPU by shape
- Storage by tier and access pattern
- Egress by path
- Managed services by SKU
- Drive forecasts using model roadmaps, reinforcement learning and inference growth, evaluation volume, and retention policy rather than trend lines.
- Run the plan-versus-reality loop by comparing planned allocations with observed fleet occupancy weekly, identifying unrecorded trades and stale allocations, and reducing variance with planning-tools and data teams.
- Qualify each incoming capacity tranche against the forecast before signature, including its shape, region, quarter, and supporting-resource envelope for storage, egress, and CPU.
- Partner with Finance and cost-efficiency teams to turn the forecast into core drivers covering at least 80% of non-accelerator spend and direct savings efforts toward future waste.
Requirements
- Experience with capacity, demand, or supply planning for large-scale technical infrastructure, such as cloud, HPC, hyperscale, or a large internal platform.
- Ability to articulate the difference between a forecast, a plan, and an allocation.
- Hands-on experience building models using SQL against a data warehouse, Python/pandas, or a forecasting and optimization stack.
- Understanding of data-center resource classes, including why storage and egress do not forecast like GPUs and why a contract's headline chip count is rarely the binding constraint.
- Preference for simple, inspectable models and the ability to instrument forecast error.
Preferred Qualifications
- Direct experience with cloud or neocloud providers involving reserved-capacity onboarding, private offers, or capacity commitments.
- Demand planning or forecasting experience for accelerator fleets, including translating research or product roadmaps into resource requirements.
- Data-center or colocation delivery experience, including power and space planning, network turn-up, site acceptance criteria, and vendor management.
- Experience with accelerator health and burn-in, collective-communications sanity testing, or fleet-health SLOs, including the ability to define health rigorously.
- Experience building lifecycle or state-machine services or systems of record for infrastructure assets.
- Experience onboarding a new hardware generation into an existing scheduler and observability stack.
Education and Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- The role follows a location-based hybrid policy. Staff are currently expected to work from one of the company's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for candidates receiving an offer, although sponsorship may not be possible for every role or candidate.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.