ML/Research Engineer, Safeguards

USD 350,000-500,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Communication @ 6 Data Pipelines Machine Learning @ 5 Python @ 5 Reinforcement Learning

Details

Anthropic is looking for ML Engineers and Research Engineers to help detect and mitigate misuse of AI systems. The Safeguards ML team builds systems that identify harmful use, including individual policy violations and sophisticated coordinated attacks, and develops defenses to keep products safe as capabilities advance. The work also includes protecting user wellbeing and ensuring models behave appropriately across a wide range of contexts, contributing to Anthropic's Responsible Scaling Policy commitments.

Responsibilities

  • Develop classifiers to detect misuse and anomalous behavior at scale.
  • Develop synthetic data pipelines for training classifiers and methods to automatically source representative evaluations.
  • Build systems to monitor harms spanning multiple exchanges, including coordinated cyber attacks and influence operations.
  • Develop methods for aggregating and analyzing signals across contexts.
  • Evaluate and improve the safety of agentic products by developing threat models and environments to test agentic risks.
  • Develop and deploy mitigations for prompt injection attacks.
  • Conduct research on automated red-teaming, adversarial robustness, and other methods for testing or finding misuse.

Requirements

  • At least 4 years of experience in ML engineering, research engineering, or applied research in academia or industry.
  • Proficiency in Python and experience building machine learning systems.
  • Ability to work across the research-to-deployment pipeline, from exploratory experiments to production systems.
  • Interest in mitigating AI misuse risks.
  • Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.
  • Bachelor's degree or an equivalent combination of education, training, and experience.
  • Relevant field of study demonstrated through coursework, training, or professional experience.

Preferred Experience

  • Language modeling and transformers.
  • Building classifiers, anomaly detection systems, or behavioral ML systems.
  • Adversarial machine learning or red-teaming.
  • Interpretability or probes.
  • Reinforcement learning.
  • High-performance, large-scale ML systems.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.

More jobs at Anthropic

Similar jobs