Safeguards Enforcement Lead, User Well-Being

USD 285,000-330,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 4 Data Analysis @ 6 Data Science GenAI Generative AI @ 4 Python @ 6 Reporting @ 4 SQL @ 6

Details

Anthropic is seeking a Safeguards Enforcement Lead for its User Well-Being team. The role manages child safety, mental health, abuse and exploitation, and age assurance enforcement workflows, including the systems and processes used to detect and respond to these harms. This is a management position involving regular exposure to explicit sexual content involving minors and potentially violent or psychologically disturbing material.

Responsibilities

  • Manage a team of individual contributors across multiple User Well-Being policy areas.
  • Serve as the primary point of contact for content review partners, including onboarding, training, quality assurance, and ongoing relationship management.
  • Design and improve enforcement workflows to scale with increasing volume while maintaining accuracy and consistency.
  • Partner with Engineering and Data Science teams to optimize detection models and automated enforcement systems.
  • Develop and maintain internal documentation, decision trees, and review guidelines for accurate and consistent enforcement.
  • Monitor emerging AI policy enforcement best practices, legal frameworks, and technology developments.
  • Identify and report misuse trends to Product, Policy, Engineering, Legal, and Trust & Safety stakeholders.
  • Coordinate reporting obligations to external bodies such as NCMEC in accordance with applicable law and Anthropic policy.

Requirements

  • Experience managing teams in the User Well-Being space.
  • Experience in trust and safety, content moderation operations, or policy enforcement, with direct exposure to child safety, mental health, abuse and exploitation, and age assurance harms.
  • Experience managing or coordinating content review operations, including quality assurance and workflow management.
  • Experience establishing and scaling policy enforcement or content review workflows.
  • Proficiency in SQL and/or other data analysis tools for monitoring workflow health, review queue metrics, and enforcement trends.
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
  • Understanding of implementing product policies at scale in content moderation.
  • Bachelor's degree or equivalent education, training, and/or experience.

Preferred Qualifications

  • Deep expertise in child safety, child sexual exploitation and abuse, online child protection, mental wellness, and age assurance.
  • Experience working with or reporting to NCMEC, IWF, or equivalent child safety reporting bodies.
  • Familiarity with CSAM reporting obligations, KOSA, COPPA, or equivalent international legal and regulatory frameworks.
  • Experience with generative AI products and AI misuse risks involving abusive content.
  • Experience designing or evaluating trauma-informed support structures and wellness protocols for content reviewers.
  • Proficiency in Python for workflow automation or data analysis.
  • Experience with hash-matching technologies such as PhotoDNA and CSAI Match, or other perceptual hashing tools used in CSAM detection.
  • Familiarity with age assurance technologies.

Benefits and Logistics

  • Annual salary of $285,000–$330,000 USD.
  • Staff are currently expected to work from an Anthropic office at least 25% of the time; some roles may require more office attendance.
  • Anthropic sponsors visas and provides immigration lawyer support, although sponsorship cannot be guaranteed for every role or candidate.
  • Wellness resources and support are provided for employees working with harmful material.
  • Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.

More jobs at Anthropic

Similar jobs