Technical Program Manager, Infrastructure

USD 290,000-365,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 AWS @ 4 Azure @ 4 CI/CD @ 4 Communication @ 7 Distributed Systems @ 6 GCP @ 4 GPU @ 4 Kubernetes @ 4 Machine Learning Observability @ 3 Security

Details

Anthropic's Infrastructure organization builds and operates the systems supporting frontier AI research and Claude, including large-scale training clusters, production infrastructure, and developer platforms. This role coordinates complex infrastructure programs across research, engineering, and product teams while supporting security, reliability, and scalability.

Responsibilities

Developer Productivity and Tooling

  • Drive cross-functional programs to improve developer environments, CI/CD infrastructure, and release processes while maintaining high security standards.
  • Coordinate large-scale migrations and platform modernization efforts across engineering teams.
  • Measure and improve developer productivity metrics, identify bottlenecks, and drive systematic improvements.
  • Lead initiatives integrating AI tools into development workflows.

Infrastructure Reliability and Operations

  • Drive programs to establish and achieve reliability targets across training infrastructure and production services.
  • Coordinate incident response improvements, post-mortem processes, and on-call rotations.
  • Establish metrics and dashboards tracking infrastructure health, capacity utilization, and operational excellence.

Cross-Functional Coordination

  • Bridge infrastructure teams, research, and product by translating technical complexities into clear updates for varied audiences.
  • Understand infrastructure, data, and compute needs and identify solutions supporting frontier research and product development.
  • Drive alignment on priorities and timelines across teams with competing constraints.

Requirements

  • 5+ years of technical program management experience delivering complex infrastructure programs in ML/AI systems or large-scale distributed systems.
  • Deep technical understanding of infrastructure systems, with the ability to engage substantively with engineers, identify technical risks, and contribute beyond project tracking.
  • Ability to create structure and processes in ambiguous environments and bring clarity to complex cross-team initiatives.
  • Strong stakeholder management and communication skills with technical and non-technical partners.
  • Ability to navigate competing priorities and use data to drive technical decisions.
  • Experience with developer productivity initiatives, CI/CD systems, or infrastructure scaling.
  • Experience with Kubernetes, cloud platforms such as AWS, GCP, or Azure, and ML infrastructure including GPU, TPU, or Trainium clusters.
  • Background working with research teams and translating their needs into concrete technical requirements.
  • Experience driving adoption of AI tools to improve engineering productivity.
  • Familiarity with observability tooling and practices.
  • Bachelor's degree or an equivalent combination of education, training, and experience in a relevant field.

Compensation

Annual salary: $290,000–$365,000 USD.

Work Arrangement and Sponsorship

The role follows a location-based hybrid policy, with staff expected to work from an Anthropic office at least 25% of the time. Anthropic sponsors visas and will make every reasonable effort to obtain a visa for successful candidates, with support from an immigration lawyer.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs