Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Agentic Systems @ 4
Data Analysis @ 7
Data Pipelines @ 4
Distributed Systems @ 4
LLM @ 4
Machine Learning @ 7
Prompt Engineering @ 4
Reinforcement Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Staff+ Software Engineer to build the methods and infrastructure used to evaluate the safety of AI models and agentic systems. The role sits at the intersection of applied machine learning research and engineering, covering evaluation methodology, dataset development, agentic investigation systems, and production tooling.
Responsibilities
- Design and run experiments to improve evaluation quality, including methods for generating representative test data, simulating realistic user behavior, and validating grading accuracy.
- Research how multi-turn conversations, tools, long context, and user diversity affect model safety behavior.
- Analyze evaluation coverage, identify measurement gaps, and evolve evaluations so they remain high-signal as model and agent capabilities advance.
- Build and own the evaluation harness for an agentic investigation system, including metrics, test cases, and grading approaches.
- Measure agent performance end-to-end, including detection precision and recall, investigation quality, and robustness.
- Drive hill-climbing on difficult harm areas and construct reinforcement learning environments to improve Claude’s safety investigation capabilities.
- Construct high-quality evaluation datasets representing real-world misuse across areas such as cyber attacks, biological weapons, and influence operations.
- Draw on real traffic patterns and synthetic generation when developing datasets.
- Collaborate with Policy and Enforcement teams to translate observed harm patterns into measurable evaluations.
- Ship successful research into evaluation, regression, and release pipelines that run during model training, after agent changes, following prompt updates, during underlying model upgrades, and beyond launch.
- Build tooling that enables policy experts to author, run, and iterate on evaluations without engineering support.
- Surface findings to research and training teams to drive upstream model improvements.
Requirements
- At least 8 years of industry software engineering or machine learning engineering experience.
- Experience building and maintaining data pipelines.
- Experience working with large language models and an understanding of their capabilities and failure modes, especially agentic systems with tool use and multi-step reasoning.
- Strong data analysis skills and the ability to draw reliable insights from large datasets.
- Ability to move between research prototyping and production-quality code.
- Ability to translate ambiguous problems into concrete, testable experiments.
- Strong interest in AI safety and its practical impact.
- Bachelor’s degree or an equivalent combination of education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Expertise in building or contributing to LLM or agent evaluation frameworks, benchmarks, or automated grading systems.
- Extensive experience in trust and safety, content moderation, or abuse detection systems.
- Experience in red teaming, adversarial testing, or jailbreak research on AI systems.
- Experience with synthetic data generation or data augmentation.
- Experience with distributed systems or large-scale data processing.
- Experience with prompt engineering or building LLM-powered applications.
Benefits
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
Work Arrangement
Anthropic expects staff to work from one of its offices at least 25% of the time. The role is listed in San Francisco, California, and New York City, New York.
Visa Sponsorship
Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for an offer recipient, with support from an immigration lawyer. Sponsorship may not be available for every role or candidate.
More jobs at Anthropic
Business Systems Analyst
Anthropic · London, United Kingdom
GBP 130,000-165,000 per year
Product Manager, Business Technology
Anthropic · London, United Kingdom
GBP 190,000-240,000 per year
Incident Manager - Detection & Response
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Staff+ Software Engineer, Auth & Identity
Anthropic · New York City, United States, San Francisco, United States
USD 405,000-485,000 per year
Accounting Analytics & BI Engineer
Anthropic · San Francisco, United States
USD 220,000-270,000 per year
Similar jobs
Research Scientist, Life Sciences
Anthropic · San Francisco, United States
USD 300,000-320,000 per year
Member of Technical Staff - Multimodal Understanding
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Staff Software Engineer, Environments Infrastructure
Anthropic · New York City, United States, San Francisco, United States
USD 405,000-605,000 per year
Research Engineer, Model Evaluations
Anthropic · New York City, United States, San Francisco, United States
USD 500,000-850,000 per year
Senior Software Engineer, Backend - Platform (Core AI Automation)
Coinbase · United States
USD 186,100-218,900 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Machine Learning Engineer, API Multicloud
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Staff Machine Learning Engineer, Consumer
Reddit · United States
USD 230,000-322,000 per year