Cyber Evaluations Engineer

USD 300,000-405,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 Fraud @ 3 Machine Learning @ 3 Python @ 5 Security @ 3

Details

Anthropic is hiring a Cyber Evaluations Engineer to build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in its AI models. The role involves designing new evaluations, conducting per-release robustness testing, analyzing jailbreak and prompt-bypass data, and developing probes to detect cyber abuse in production. The engineer will also help shape the overall abuse-detection architecture in partnership with the policy team.

Responsibilities

  • Design and run capability, uplift, and safety evaluations to assess cyber-relevant risks in new models.
  • Execute safeguard-robustness testing ahead of major model launches.
  • Analyze evaluation results and communicate findings clearly to teams and stakeholders.
  • Design, prototype, and tune detection probes for cyber misuse.
  • Work with the cyber policy team to translate policy lines into a layered, robust abuse-detection architecture.
  • Measure detection precision and coverage over time.
  • Build and maintain internal tooling used to run and score evaluations.
  • Collaborate with policy and engineering partners to translate evaluation findings into safeguard improvements.

Requirements

  • Experience building or running evaluations, benchmarks, or test suites for software or machine-learning systems, including delivering results on short, fixed timelines.
  • Hands-on cybersecurity experience, such as CTF participation, vulnerability research, exploit development, or security research.
  • Proficiency in Python.
  • Strong ability to communicate evaluation results with multiple cross-functional stakeholders or policy stakeholders.
  • A bachelor's degree or an equivalent combination of education, training, and experience.
  • A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Deep offensive-security or security-research experience, including experience building AI security benchmarks.
  • Experience analyzing adversarial or abuse data, such as jailbreaks, prompt bypasses, intrusion telemetry, or fraud telemetry.
  • Experience working onsite with government partners on testing or evaluation engagements.
  • Experience with AI/ML evaluation frameworks.
  • Familiarity with coordinated vulnerability disclosure practices.
  • Experience testing pre-release or pre-deployment software or models under confidentiality constraints.
  • Experience authoring detection content, such as Sigma, YARA, Suricata, or SIEM rules, or building ML-based abuse detection.
  • An active secret security clearance or higher, or eligibility to obtain one.

Compensation

  • Annual salary: $300,000–$405,000 USD.

Logistics

  • Location-based hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
  • Anthropic sponsors visas, though sponsorship is not guaranteed for every role or candidate. The company retains an immigration lawyer to assist with visa processes.
  • Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.

More jobs at Anthropic

Similar jobs