ML/Research Engineer, Safeguards

USD 350,000-500,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Communication @ 6 Data Pipelines Machine Learning Python @ 5 Reinforcement Learning

Details

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. This role helps build safeguards for Anthropic’s AI systems to identify and mitigate misuse.

Responsibilities

  • Develop classifiers to detect misuse and anomalous behavior at scale. This includes developing synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on.
  • Build systems to monitor for harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts.
  • Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks.
  • Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse.

Requirements

  • 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry.
  • Proficiency in Python and experience building ML systems.
  • Comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems.
  • Concerned about misuse risks of AI systems and wants to work to mitigate them.
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.

Strong candidates may also have experience with

  • Language modeling and transformers.
  • Building classifiers, anomaly detection systems, or behavioral ML.
  • Adversarial machine learning or red-teaming.
  • Interpretability or probes.
  • Reinforcement learning.
  • High-performance, large-scale ML systems.

Logistics

  • Annual Salary: $350,000 - $500,000 USD.
  • Location-based hybrid policy: expect all staff to be in one of Anthropic’s offices at least 25% of the time.
  • Visa sponsorship: Anthropic sponsors visas, and if they make an offer they will make every reasonable effort to get you a visa and retain an immigration lawyer to help with this.

More jobs at Anthropic

Similar jobs