Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Claude Code @ 3
Communication @ 3
SQL @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Safeguards team enforces policies, protects users, and helps ensure that its platform is not misused. This role focuses on safety evaluations and supports model launch readiness by running and monitoring evaluations, interpreting results, driving mitigations, coordinating the creation of new evaluations, and building scalable processes and documentation.
The role is highly cross-functional and involves partnering with policy experts, Safeguards engineering teams, product teams, and other stakeholders to ensure evaluations remain comprehensive and current as policies, threat vectors, model capabilities, and product surfaces evolve.
Responsibilities
- Support model launch readiness by running evaluations, monitoring and interpreting results, and surfacing regressions or unexpected behavior changes.
- Partner with policy and domain experts across the evaluation lifecycle, including risk identification, evaluation scoping, creation of new evaluations, and maintenance of existing evaluations.
- Help manage evaluation outcomes, interpret results, and drive mitigations with cross-functional stakeholders.
- Develop processes and evaluation paradigms that remain high-signal and insightful as models improve.
- Build processes and frameworks for product-specific evaluations as Anthropic's product surface expands.
- Help design and scope tooling improvements that support evolving evaluation needs and enable self-service evaluation creation and iteration for non-technical users.
- Write and maintain documentation for evaluation creation, execution, and interpretation.
Requirements
- Experience in trust and safety, content operations, policy enforcement, or a related operational role at a technology company.
- Ability to work effectively in ambiguous, fast-moving environments.
- Experience building processes, workflows, or programs from scratch.
- Strong program management skills, including tracking timelines, dependencies, and deliverables across complex, multi-stakeholder efforts.
- Interest in expanding technical skills through internal tools and AI-assisted workflows, such as Claude Code.
- Ability to manage multiple concurrent workstreams, prioritize effectively, and switch contexts while maintaining attention to detail.
- Strong generalist skills and sound judgment when working with incomplete information.
- Clear and concise written and cross-functional communication.
- Bachelor's degree or an equivalent combination of education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Strong candidates may also have experience with high-stakes timelines such as product launches, incident response, or regulatory deadlines; coordinating across engineering, policy, and product teams; developing SOPs, runbooks, and operational documentation; and using data tools such as SQL, dashboards, and spreadsheets. Comfort working with sensitive content areas is also beneficial.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from an office at least 25% of the time, although some roles may require more frequent office attendance.