Technical Program Manager, Safeguards (Infrastructure & Evals)

USD 290,000-365,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 Datadog @ 2 SRE @ 3

Details

Anthropic’s Safeguards Engineering team builds and operates infrastructure that keeps AI systems safe in production, including classifiers, detection pipelines, evaluation platforms, and monitoring systems.

This Technical Program Manager will own the operational health and forward momentum of the Safeguards Infrastructure and Evals stack. The role focuses on reliability, incident response, post-mortems, service-level objectives, runbooks, infrastructure migrations, evaluation-platform improvements, and cross-team dependencies. The position requires sufficient technical depth to triage incidents, assess safety-critical issues, and work effectively with engineering teams, while focusing primarily on operational execution and program management.

Responsibilities

  • Own the Safeguards Engineering operations review cadence, including incident and failure reviews, reliability trends, and operational-risk decisions.
  • Drive incident tracking and post-mortem execution across the organization, including incidents owned by partner teams such as Inference. Ensure post-mortems are completed and action items are closed.
  • Establish and maintain service-level objectives with Safeguards Engineering, Inference, and Cloud Inference teams, including tracking and reporting on SLO performance.
  • Maintain the quality, accuracy, and actionability of runbooks and ensure incident ownership is clear.
  • Manage infrastructure projects and migrations, including migrations between infrastructure platforms, incident platforms, and cloud-system monitoring solutions.
  • Coordinate evaluation-platform improvements with the Evals Engineering team, including self-serve capabilities and broader evaluation-factory infrastructure.
  • Coordinate cross-team dependencies, sequence work, track progress, and maintain momentum across operational and infrastructure initiatives.

Requirements

  • Solid technical program management experience, particularly in operational or infrastructure-heavy environments.
  • Understanding of production machine-learning systems sufficient to triage incidents and have substantive technical discussions with engineers.
  • Ability to establish processes and follow-ups that close post-mortem actions, maintain SLOs, and keep runbooks current.
  • Ability to coordinate effectively across team boundaries and influence partner teams without direct authority.
  • Ability to manage both operational work and longer-term platform projects while context-switching effectively.
  • Experience with or strong interest in AI safety and the reliability requirements of safety-critical pipelines.
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Experience with SRE practices, incident-management frameworks, or large-scale on-call operations.
  • Experience with evaluation infrastructure for machine-learning systems.
  • Experience driving infrastructure migrations in complex, multi-team environments where operational systems cannot go offline.
  • Familiarity with monitoring and alerting tools such as PagerDuty, Datadog, or equivalents.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

The role has a rolling application process with no specified deadline.

More jobs at Anthropic

Similar jobs