Staff+ Software Engineer, Safeguards Review Tooling

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 API Agentic Systems @ 4 Audit Communication @ 6 Compliance @ 4 Data Science Fraud @ 4 LLM

Details

Anthropic is seeking engineers for its Safeguards Review Tooling team, which builds systems used by human reviewers and Claude to investigate potential harms and take enforcement actions across first-party products and third-party cloud platforms. The role involves owning investigation tools and the underlying platform, including analytics capabilities, privacy-preserving primitives, and sandbox environments for developing review interfaces and workflows.

Responsibilities

  • Build investigation, review, and enforcement tooling, including case queues, investigation views, decision and audit logging, and account-actioning workflows.
  • Develop reusable APIs, data storage, and backend services for rapidly and safely launching review workflows.
  • Scale review through automation, including Claude-assisted and Claude-driven review workflows while keeping humans involved where their judgment is important.
  • Partner with policy, operations, legal, privacy, and data science teams to translate investigation and enforcement needs into reliable systems.
  • Implement granular permissions, audit trails, data-access controls, content obfuscation, and exposure limits.
  • Instrument tools with metrics for queue health, reviewer throughput, and decision quality.
  • Evolve tooling alongside privacy primitives and data-retention commitments.

Requirements

  • Technical background in full-stack or platform engineering, with the ability to engage deeply in architecture and design discussions.
  • Experience shipping internal tools or platforms for demanding operational users and measurably improving their workflows.
  • Experience working cross-functionally with operations, policy, legal, or other non-engineering partners.
  • Excellent communication skills and the ability to explain technical tradeoffs to non-technical stakeholders.
  • Interest in the societal impacts of AI and in making powerful systems safer.
  • 8+ years of industry software engineering experience preferred.
  • Experience with trust and safety, integrity, fraud, abuse-prevention, or human-review tooling preferred.
  • Experience designing systems under strict privacy, compliance, or data-governance constraints preferred.
  • Experience integrating LLMs or agentic systems into operational workflows or building human-in-the-loop automation preferred.
  • Experience building developer platforms or extensible tooling frameworks preferred.
  • Experience supporting enforcement or moderation systems across multiple product surfaces, including enterprise or cloud platforms, preferred.
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.

Compensation

Annual salary: $320,000–$485,000 USD.

Work Policy and Sponsorship

Anthropic currently expects staff to work from one of its offices at least 25% of the time. Anthropic explicitly states that it sponsors visas and will make reasonable efforts to obtain a visa for successful candidates, with support from an immigration lawyer.

More jobs at Anthropic

Similar jobs