Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Azure @ 4
Communication @ 7
GCP @ 4
GPU @ 4
HPC @ 4
Kubernetes @ 3
Leadership @ 7
Machine Learning @ 4
Observability @ 4
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Technical Program Manager on the Compute team, you will drive the planning, coordination, and execution of programs that keep Anthropic's compute infrastructure running efficiently at scale. You will own critical workstreams across the compute lifecycle, including supply procurement, capacity onboarding, allocation, utilization, workload migrations, and decommissions.
You will partner with Infrastructure, Systems, Research, Finance, and Capacity Engineering teams to develop the processes, tooling, and coordination mechanisms needed to manage a complex compute environment.
Responsibilities
- Own and drive critical programs across the compute lifecycle, coordinating execution across engineering, research, and operations teams.
- Build and maintain operational visibility into the compute fleet, including supply, demand, utilization, and health.
- Lead cross-functional coordination for compute transitions, including bringing new capacity online, migrating workloads, and managing decommissions across cloud providers and hardware platforms.
- Partner with engineering and research leadership to resolve competing priorities and align compute planning, allocation, and usage.
- Identify and close operational gaps through new tooling, improved processes, and stronger cross-team communication.
- Lead trade-off discussions involving utilization, cost, latency, and reliability, and communicate decisions to leadership.
- Develop and improve frameworks for planning, tracking, and executing compute programs at increasing scale and complexity.
Requirements
- 7+ years of technical program management experience in infrastructure, platform engineering, or compute-intensive environments.
- Experience leading complex, cross-functional programs involving multiple engineering teams, competing priorities, and ambiguous requirements.
- Experience working with research or machine learning teams and translating their needs into operational plans and technical requirements.
- Ability to understand technical details involving cloud infrastructure, cluster management, job scheduling, and resource orchestration while maintaining program-level visibility.
- Ability to define scope and build processes in ambiguous, fast-moving environments.
- Strong communication skills and the ability to work with engineers, researchers, finance teams, and executive leadership.
- Track record of building trust with engineering teams and driving change through influence rather than authority.
- Bachelor's degree or an equivalent combination of education, training, and experience in a relevant field or demonstrated through relevant coursework, training, or professional experience.
Preferred Qualifications
- Experience managing compute capacity across AWS, GCP, Azure, hybrid cloud, or on-premises environments.
- Familiarity with job scheduling, resource orchestration, or workload management systems such as Kubernetes, Slurm, Borg, YARN, or custom schedulers.
- Experience with GPU or accelerator infrastructure and large-scale machine learning training and inference workloads.
- Experience building or improving infrastructure observability, including dashboards, alerting, efficiency metrics, and cost attribution.
- Capacity planning experience involving demand forecasting, cost modeling, or hardware lifecycle management.
- Experience scaling AI/ML, HPC, or large-scale cloud environments through hypergrowth.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration. Staff are currently expected to work from one of the company's offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas and makes reasonable efforts to support visa applications, with assistance from an immigration lawyer.