Safeguards Enforcement Analyst, Safety Evaluations

USD 230,000-270,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Claude Code @ 3 Communication @ 3 SQL @ 6

Details

Anthropic's Safeguards team enforces policies, protects users, and helps ensure that its platform is not misused. This role focuses on safety evaluations and supports model launch readiness by running and monitoring evaluations, interpreting results, driving mitigations, coordinating the creation of new evaluations, and building scalable processes and documentation.

The role is highly cross-functional and involves partnering with policy experts, Safeguards engineering teams, product teams, and other stakeholders to ensure evaluations remain comprehensive and current as policies, threat vectors, model capabilities, and product surfaces evolve.

Responsibilities

  • Support model launch readiness by running evaluations, monitoring and interpreting results, and surfacing regressions or unexpected behavior changes.
  • Partner with policy and domain experts across the evaluation lifecycle, including risk identification, evaluation scoping, creation of new evaluations, and maintenance of existing evaluations.
  • Help manage evaluation outcomes, interpret results, and drive mitigations with cross-functional stakeholders.
  • Develop processes and evaluation paradigms that remain high-signal and insightful as models improve.
  • Build processes and frameworks for product-specific evaluations as Anthropic's product surface expands.
  • Help design and scope tooling improvements that support evolving evaluation needs and enable self-service evaluation creation and iteration for non-technical users.
  • Write and maintain documentation for evaluation creation, execution, and interpretation.

Requirements

  • Experience in trust and safety, content operations, policy enforcement, or a related operational role at a technology company.
  • Ability to work effectively in ambiguous, fast-moving environments.
  • Experience building processes, workflows, or programs from scratch.
  • Strong program management skills, including tracking timelines, dependencies, and deliverables across complex, multi-stakeholder efforts.
  • Interest in expanding technical skills through internal tools and AI-assisted workflows, such as Claude Code.
  • Ability to manage multiple concurrent workstreams, prioritize effectively, and switch contexts while maintaining attention to detail.
  • Strong generalist skills and sound judgment when working with incomplete information.
  • Clear and concise written and cross-functional communication.
  • Bachelor's degree or an equivalent combination of education, training, and experience.
  • Relevant field of study demonstrated through coursework, training, or professional experience.

Strong candidates may also have experience with high-stakes timelines such as product launches, incident response, or regulatory deadlines; coordinating across engineering, policy, and product teams; developing SOPs, runbooks, and operational documentation; and using data tools such as SQL, dashboards, and spreadsheets. Comfort working with sensitive content areas is also beneficial.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from an office at least 25% of the time, although some roles may require more frequent office attendance.

More jobs at Anthropic

Similar jobs