Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Fraud @ 3
Machine Learning @ 3
Python @ 5
Security @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is hiring a Cyber Evaluations Engineer to build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in its AI models. The role involves designing new evaluations, conducting per-release robustness testing, analyzing jailbreak and prompt-bypass data, and developing probes to detect cyber abuse in production. The engineer will also help shape the overall abuse-detection architecture in partnership with the policy team.
Responsibilities
- Design and run capability, uplift, and safety evaluations to assess cyber-relevant risks in new models.
- Execute safeguard-robustness testing ahead of major model launches.
- Analyze evaluation results and communicate findings clearly to teams and stakeholders.
- Design, prototype, and tune detection probes for cyber misuse.
- Work with the cyber policy team to translate policy lines into a layered, robust abuse-detection architecture.
- Measure detection precision and coverage over time.
- Build and maintain internal tooling used to run and score evaluations.
- Collaborate with policy and engineering partners to translate evaluation findings into safeguard improvements.
Requirements
- Experience building or running evaluations, benchmarks, or test suites for software or machine-learning systems, including delivering results on short, fixed timelines.
- Hands-on cybersecurity experience, such as CTF participation, vulnerability research, exploit development, or security research.
- Proficiency in Python.
- Strong ability to communicate evaluation results with multiple cross-functional stakeholders or policy stakeholders.
- A bachelor's degree or an equivalent combination of education, training, and experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Deep offensive-security or security-research experience, including experience building AI security benchmarks.
- Experience analyzing adversarial or abuse data, such as jailbreaks, prompt bypasses, intrusion telemetry, or fraud telemetry.
- Experience working onsite with government partners on testing or evaluation engagements.
- Experience with AI/ML evaluation frameworks.
- Familiarity with coordinated vulnerability disclosure practices.
- Experience testing pre-release or pre-deployment software or models under confidentiality constraints.
- Experience authoring detection content, such as Sigma, YARA, Suricata, or SIEM rules, or building ML-based abuse detection.
- An active secret security clearance or higher, or eligibility to obtain one.
Compensation
- Annual salary: $300,000–$405,000 USD.
Logistics
- Location-based hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas, though sponsorship is not guaranteed for every role or candidate. The company retains an immigration lawyer to assist with visa processes.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.
More jobs at Anthropic
Staff+ Software Engineer, Distributed Systems
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-485,000 per year
Salesforce Developer, Partnerships
Anthropic · New York City, United States, San Francisco, United States
USD 270,000-345,000 per year
Forward Deployed Engineer
Anthropic · London, United Kingdom
GBP 225,000-255,000 per year
Recruiting Analytics Data Engineer
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 285,000-380,000 per year
AI Deployment Specialist, Beneficial Deployments
Anthropic · New York City, United States, San Francisco, United States
USD 0 per year
Similar jobs
Staff+ Software Engineer, Account Abuse
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Data Scientist, Safety
OpenAI · New York City, United States, San Francisco, United States
USD 230,000-325,000 per year
Staff Machine Learning Engineer, AI Security
Reddit · United States
USD 230,000-322,000 per year
Lead, Security Controls Assurance - SOX
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 410,000-510,000 per year
Staff+ Software Engineer, Platform Portability
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Data Scientist, Cybersecurity
OpenAI · United States, New York City, United States, San Francisco, United States
USD 263,000-515,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Security Engineer, Offensive Security
Anthropic · New York City, United States, San Francisco, United States
USD 300,000-320,000 per year