Staff+ Software Engineer, Safeguards Evals

USD 320,000-485,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 7 Agentic Systems @ 4 Data Analysis @ 7 Data Pipelines @ 4 Distributed Systems @ 4 LLM @ 4 Machine Learning @ 7 Prompt Engineering @ 4 Reinforcement Learning

Details

Anthropic is seeking a Staff+ Software Engineer to build the methods and infrastructure used to evaluate the safety of AI models and agentic systems. The role sits at the intersection of applied machine learning research and engineering, covering evaluation methodology, dataset development, agentic investigation systems, and production tooling.

Responsibilities

  • Design and run experiments to improve evaluation quality, including methods for generating representative test data, simulating realistic user behavior, and validating grading accuracy.
  • Research how multi-turn conversations, tools, long context, and user diversity affect model safety behavior.
  • Analyze evaluation coverage, identify measurement gaps, and evolve evaluations so they remain high-signal as model and agent capabilities advance.
  • Build and own the evaluation harness for an agentic investigation system, including metrics, test cases, and grading approaches.
  • Measure agent performance end-to-end, including detection precision and recall, investigation quality, and robustness.
  • Drive hill-climbing on difficult harm areas and construct reinforcement learning environments to improve Claude’s safety investigation capabilities.
  • Construct high-quality evaluation datasets representing real-world misuse across areas such as cyber attacks, biological weapons, and influence operations.
  • Draw on real traffic patterns and synthetic generation when developing datasets.
  • Collaborate with Policy and Enforcement teams to translate observed harm patterns into measurable evaluations.
  • Ship successful research into evaluation, regression, and release pipelines that run during model training, after agent changes, following prompt updates, during underlying model upgrades, and beyond launch.
  • Build tooling that enables policy experts to author, run, and iterate on evaluations without engineering support.
  • Surface findings to research and training teams to drive upstream model improvements.

Requirements

  • At least 8 years of industry software engineering or machine learning engineering experience.
  • Experience building and maintaining data pipelines.
  • Experience working with large language models and an understanding of their capabilities and failure modes, especially agentic systems with tool use and multi-step reasoning.
  • Strong data analysis skills and the ability to draw reliable insights from large datasets.
  • Ability to move between research prototyping and production-quality code.
  • Ability to translate ambiguous problems into concrete, testable experiments.
  • Strong interest in AI safety and its practical impact.
  • Bachelor’s degree or an equivalent combination of education, training, and experience.
  • Relevant field of study demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Expertise in building or contributing to LLM or agent evaluation frameworks, benchmarks, or automated grading systems.
  • Extensive experience in trust and safety, content moderation, or abuse detection systems.
  • Experience in red teaming, adversarial testing, or jailbreak research on AI systems.
  • Experience with synthetic data generation or data augmentation.
  • Experience with distributed systems or large-scale data processing.
  • Experience with prompt engineering or building LLM-powered applications.

Benefits

  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Office space for collaboration.

Work Arrangement

Anthropic expects staff to work from one of its offices at least 25% of the time. The role is listed in San Francisco, California, and New York City, New York.

Visa Sponsorship

Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for an offer recipient, with support from an immigration lawyer. Sponsorship may not be available for every role or candidate.

More jobs at Anthropic

Similar jobs