Safeguards Enforcement Analyst, Cyber Harm

USD 285,000-330,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Data Analysis @ 5 Data Science GenAI Generative AI @ 3 LLM Python @ 5 SQL @ 5

Details

Review content and execute enforcement actions across Anthropic's products and services, focusing on detecting and mitigating attempts to misuse AI systems for malicious cyber operations. The initial focus is on flagged activity related to cyberattacks, malware development, offensive exploitation, and other harmful cyber operations. The role may later expand to broader enforcement areas and may involve exposure to explicit, violent, technical, or psychologically disturbing content. Weekend and holiday escalation response may be required.

Responsibilities

  • Review flagged content and accounts and make accurate, well-documented enforcement decisions in line with usage policies.
  • Detect and mitigate potential misuse of AI systems to facilitate cyberattacks, malware creation, exploitation tooling, and related harmful cyber operations.
  • Triage and escalate novel, ambiguous, or high-severity cases to appropriate stakeholders.
  • Provide detailed feedback to the Safeguards policy design team about policy gaps identified through enforcement scenarios.
  • Partner with Engineering and Data Science teams by surfacing detection model errors and quality signals from review to improve precision and recall.
  • Maintain high accuracy and consistency across review queues.
  • Stay current on AI policy enforcement best practices, threat actor tactics, and the evolving cyber threat landscape.

Requirements

  • Experience in cybersecurity, including knowledge of offensive techniques, exploit development, malware analysis, or vulnerability research.
  • Experience performing content review, abuse investigations, or policy enforcement at volume.
  • Proficiency in SQL and/or Python for data analysis and threat detection.
  • Experience identifying emerging risks and communicating findings to Product, Policy, Engineering, and Legal stakeholders.
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement.
  • Bachelor's degree or an equivalent combination of education, training, and experience.
  • A relevant field of study demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Experience in trust and safety, abuse investigations, cybersecurity investigations, or threat intelligence at a technology or AI company.
  • Experience with large language models and understanding of how AI technology can be misused for cyber operations.
  • Experience operating abuse monitoring programs or enforcement review systems.
  • Understanding of implementing product policies at scale, including content moderation.
  • Experience working with government agencies, regulated environments, or information-sharing communities.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.

More jobs at Anthropic

Similar jobs