Safeguards Enforcement Lead, Cyber Harms

USD 285,000-330,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 4 Data Analysis @ 6 Data Science GenAI Generative AI @ 4 LLM Python @ 6 SQL @ 6

Details

Manage and execute enforcement actions across Anthropic’s products and services, focusing on detecting and mitigating attempts to misuse AI systems for malicious cyber operations. Develop strategic enforcement frameworks for activity related to cyberattacks, malware development, offensive exploitation, and other harmful cyber operations. Manage a team of Cyber Enforcement Analysts and contractors implementing the enforcement strategy.

The role may involve exposure to explicit content, including violent, technical, or psychologically disturbing material, and may require responding to escalations during weekends and holidays.

Responsibilities

  • Manage a team of Cyber Enforcement Analysts and contractors and oversee the vision for Cyber Enforcement strategy.
  • Create strategies to detect and mitigate potential misuse of AI systems to facilitate cyberattacks, malware creation, exploitation tooling, and related harmful cyber operations.
  • Collaborate with stakeholders on novel, ambiguous, or high-severity cases.
  • Collaborate with the Safeguards Policy Design Team on policy gaps identified through real enforcement scenarios.
  • Partner with Engineering and Data Science teams to ensure tooling and measurement support enforcement operations.
  • Keep up to date with AI policy enforcement best practices, threat actor tactics, and the evolving cyber threat landscape, using this knowledge to inform enforcement decisions.

Requirements

  • Experience as a people manager.
  • Experience in cybersecurity, including knowledge of offensive techniques, exploit development, malware analysis, or vulnerability research.
  • Experience performing content review, abuse investigations, or policy enforcement at volume.
  • Proficiency in SQL and/or Python for data analysis and threat detection.
  • Experience identifying emerging risks and communicating findings to diverse stakeholders, including Product, Policy, Engineering, and Legal teams.
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement.
  • Bachelor’s degree or an equivalent combination of education, training, and experience.
  • A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Experience in trust and safety, abuse investigations, cybersecurity investigations, or threat intelligence within a technology or AI company.
  • Experience with large language models and an understanding of how AI technology could be misused for cyber operations.
  • Experience operating within abuse monitoring programs or enforcement review systems.
  • Understanding of the challenges involved in implementing product policies at scale, including in content moderation.
  • Experience working with government agencies, regulated environments, or information-sharing communities.

Compensation

Annual salary: $285,000–$330,000 USD.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. The company sponsors visas and retains an immigration lawyer to assist with visa applications, although sponsorship availability may vary by role and candidate.

Work Arrangement

Remote-friendly with travel required. Staff are currently expected to be in one of Anthropic’s offices at least 25% of the time, although some roles may require more office time.

More jobs at Anthropic

Similar jobs