Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 6
Data Pipelines
Machine Learning
Python @ 5
Reinforcement Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. This role helps build safeguards for Anthropic’s AI systems to identify and mitigate misuse.
Responsibilities
- Develop classifiers to detect misuse and anomalous behavior at scale. This includes developing synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on.
- Build systems to monitor for harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts.
- Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks.
- Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse.
Requirements
- 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry.
- Proficiency in Python and experience building ML systems.
- Comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems.
- Concerned about misuse risks of AI systems and wants to work to mitigate them.
- Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.
Strong candidates may also have experience with
- Language modeling and transformers.
- Building classifiers, anomaly detection systems, or behavioral ML.
- Adversarial machine learning or red-teaming.
- Interpretability or probes.
- Reinforcement learning.
- High-performance, large-scale ML systems.
Logistics
- Annual Salary: $350,000 - $500,000 USD.
- Location-based hybrid policy: expect all staff to be in one of Anthropic’s offices at least 25% of the time.
- Visa sponsorship: Anthropic sponsors visas, and if they make an offer they will make every reasonable effort to get you a visa and retain an immigration lawyer to help with this.
More jobs at Anthropic
Recruiter, Applied AI
Anthropic · San Francisco, United States
USD 175,000-240,000 per year
Staff Software Engineer, GTM Systems
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Evals Infrastructure Tech Lead / Manager
Anthropic · San Francisco, United States
USD 500,000-850,000 per year
Manager, IT Support
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-265,000 per year
Technical Program Manager, GTM Systems
Anthropic · San Francisco, United States, New York City, United States
USD 290,000-365,000 per year
Similar jobs
Research Engineer, Life Sciences
Anthropic · San Francisco, United States
USD 350,000-500,000 per year
Security Software Engineer, Detection & Response Platform
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Research Engineer, Computer Use
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 500,000-850,000 per year
Staff Full Stack Engineer, Identity
Stripe · South San Francisco, United States, New York City, United States, Seattle, United States
USD 224,000-336,000 per year
Research Engineer/Research Scientist, Pre-Training
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 350,000-850,000 per year
Machine Learning Infrastructure Engineer, Safeguards Research
Anthropic · San Francisco, United States, New York City, United States
USD 350,000-500,000 per year
Finance Systems Integration Engineer
Anthropic · San Francisco, United States, Seattle, United States
USD 205,000-270,000 per year
Data Operations Manager, Human Data
Anthropic · San Francisco, United States, New York City, United States
USD 270,000-365,000 per year