Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 6
Data Pipelines
Machine Learning @ 5
Python @ 5
Reinforcement Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is looking for ML Engineers and Research Engineers to help detect and mitigate misuse of AI systems. The Safeguards ML team builds systems that identify harmful use, including individual policy violations and sophisticated coordinated attacks, and develops defenses to keep products safe as capabilities advance. The work also includes protecting user wellbeing and ensuring models behave appropriately across a wide range of contexts, contributing to Anthropic's Responsible Scaling Policy commitments.
Responsibilities
- Develop classifiers to detect misuse and anomalous behavior at scale.
- Develop synthetic data pipelines for training classifiers and methods to automatically source representative evaluations.
- Build systems to monitor harms spanning multiple exchanges, including coordinated cyber attacks and influence operations.
- Develop methods for aggregating and analyzing signals across contexts.
- Evaluate and improve the safety of agentic products by developing threat models and environments to test agentic risks.
- Develop and deploy mitigations for prompt injection attacks.
- Conduct research on automated red-teaming, adversarial robustness, and other methods for testing or finding misuse.
Requirements
- At least 4 years of experience in ML engineering, research engineering, or applied research in academia or industry.
- Proficiency in Python and experience building machine learning systems.
- Ability to work across the research-to-deployment pipeline, from exploratory experiments to production systems.
- Interest in mitigating AI misuse risks.
- Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.
- Bachelor's degree or an equivalent combination of education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Preferred Experience
- Language modeling and transformers.
- Building classifiers, anomaly detection systems, or behavioral ML systems.
- Adversarial machine learning or red-teaming.
- Interpretability or probes.
- Reinforcement learning.
- High-performance, large-scale ML systems.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.