Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Fraud @ 3
Machine Learning @ 3
Python @ 5
Security @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is hiring a Cyber Evaluations Engineer to build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in its AI models. The role involves designing new evaluations, conducting per-release robustness testing, analyzing jailbreak and prompt-bypass data, and developing probes to detect cyber abuse in production. The engineer will also help shape the overall abuse-detection architecture in partnership with the policy team.
Responsibilities
- Design and run capability, uplift, and safety evaluations to assess cyber-relevant risks in new models.
- Execute safeguard-robustness testing ahead of major model launches.
- Analyze evaluation results and communicate findings clearly to teams and stakeholders.
- Design, prototype, and tune detection probes for cyber misuse.
- Work with the cyber policy team to translate policy lines into a layered, robust abuse-detection architecture.
- Measure detection precision and coverage over time.
- Build and maintain internal tooling used to run and score evaluations.
- Collaborate with policy and engineering partners to translate evaluation findings into safeguard improvements.
Requirements
- Experience building or running evaluations, benchmarks, or test suites for software or machine-learning systems, including delivering results on short, fixed timelines.
- Hands-on cybersecurity experience, such as CTF participation, vulnerability research, exploit development, or security research.
- Proficiency in Python.
- Strong ability to communicate evaluation results with multiple cross-functional stakeholders or policy stakeholders.
- A bachelor's degree or an equivalent combination of education, training, and experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Deep offensive-security or security-research experience, including experience building AI security benchmarks.
- Experience analyzing adversarial or abuse data, such as jailbreaks, prompt bypasses, intrusion telemetry, or fraud telemetry.
- Experience working onsite with government partners on testing or evaluation engagements.
- Experience with AI/ML evaluation frameworks.
- Familiarity with coordinated vulnerability disclosure practices.
- Experience testing pre-release or pre-deployment software or models under confidentiality constraints.
- Experience authoring detection content, such as Sigma, YARA, Suricata, or SIEM rules, or building ML-based abuse detection.
- An active secret security clearance or higher, or eligibility to obtain one.
Compensation
- Annual salary: $300,000–$405,000 USD.
Logistics
- Location-based hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas, though sponsorship is not guaranteed for every role or candidate. The company retains an immigration lawyer to assist with visa processes.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.
More jobs at Anthropic
Research Manager, Biological Safety
Anthropic · San Francisco, United States
USD 405,000-485,000 per year
Senior Manager, IT SOX
Anthropic · San Francisco, United States
USD 230,000-300,000 per year
Technical Program Manager, Enterprise Readiness
Anthropic · New York City, United States, San Francisco, United States
USD 365,000-435,000 per year
Staff Software Engineer, Observability & Profiling
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Safeguards Enforcement Lead, User Well-Being
Anthropic · Washington, United States, New York City, United States, San Francisco, United States
USD 285,000-330,000 per year
Similar jobs
Staff+ Software Engineer, Account Abuse
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Data Scientist, Safety
OpenAI · New York City, United States, San Francisco, United States
USD 230,000-325,000 per year
Staff+ Software Engineer, Platform Portability
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Data Scientist, Cybersecurity
OpenAI · United States, New York City, United States, San Francisco, United States
USD 263,000-515,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Staff+ Software Engineer, Safeguards Data
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-485,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Staff Software Engineer - Site Defense
Reddit · United States
USD 217,000-303,900 per year