Head of Policy Design, Societal Harms

USD 330,000-395,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 4 API @ 4 Agentic AI @ 4 Data Science GenAI Generative AI @ 4 LLM @ 4 Reporting @ 9

Details

Anthropic's Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. This role leads the policy design team responsible for radicalization, child safety, user well-being, harmful manipulation, election integrity, and other consumer harm areas.

The role involves defining risks associated with engaging with Claude, determining how those risks materialize in the real world, and developing mitigations in partnership with research, product, and engineering. Mitigations span model training, policies, detection and enforcement systems, and product interventions. The role also includes exposure to explicit sexual, violent, or psychologically disturbing content.

Responsibilities

  • Lead, develop, and grow managers and teams responsible for the consumer harms portfolio, including child safety, user well-being, harmful manipulation, and election integrity.
  • Coordinate policy decisions across the portfolio and build mechanisms to keep decisions tracked, consistent, and understandable.
  • Set the strategy for how policies, detection and enforcement systems, and product interventions complement model training, working closely with the alignment training team.
  • Prioritize competing harm areas and communicate tradeoffs and rationale to leadership.
  • Serve as the escalation point for high-severity and ambiguous consumer harm decisions, including emerging risks.
  • Partner with engineering, data science, product, legal, and research throughout the model development cycle, from training through launch.
  • Engage external experts, civil society organizations, and regulators, translating that engagement into stronger policy and enforcement.

Requirements

  • Experience leading teams, including managing managers or senior specialists, in AI safety, product policy, or a related field.
  • Deep applied familiarity with consumer harm areas such as child safety, mental health and well-being, manipulation, or election integrity.
  • Exceptional cross-team collaboration skills and experience reaching shared decisions with teams outside direct reporting lines.
  • Working understanding of frontier model development and deployment, including training and fine-tuning, evaluations, launch processes, and the differences between consumer products, APIs, and agentic tools.
  • Experience translating policy positions into mechanisms that can be enforced and measured, and communicating reasoning to technical and non-technical audiences, including executives.
  • Sound judgment in ambiguous, high-consequence situations.

Preferred Qualifications

  • Subject-matter depth in one or more relevant harm areas through academia, clinical practice, civil society, government, or trust and safety work.
  • Experience working with model training or research teams on model behavior or the character of a deployed AI system.
  • Experience with generative AI safety systems, including LLM-based classification, evaluation, or enforcement pipelines.
  • Experience engaging child safety organizations, election authorities, mental health experts, regulators, or other external stakeholders.
  • Experience using agentic AI tools to scale team analysis and operations.

A bachelor's degree or equivalent combination of education, training, and experience is required. The field of study should be relevant to the role through coursework, training, or professional experience.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs