Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Data Analysis @ 5
Hiring @ 3
LLM
Machine Learning @ 3
People Management @ 3
Prioritization @ 3
Python @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Safeguards organization builds the policies, evaluations, and enforcement systems that keep its models from contributing to catastrophic harm. This role leads the research engineering team responsible for biological safety evaluations, datasets, and classifiers that govern how models handle biological knowledge.
The role involves managing research scientists and engineers who design and run capability evaluations against frontier models, curate training data for safety classifiers, train and iterate on classifiers with ML engineers, and measure performance against adversarial pressure in production traffic. The manager sets technical direction, determines team priorities, and owns results. This is a hands-on management role requiring sufficient technical depth to review evaluation designs, analyze classifier failure modes, and represent the work to Research, Product, and Policy partners.
The team focuses on balancing safeguards that are robust against sophisticated actors with low false-positive rates for legitimate researchers using Claude in life sciences.
Responsibilities
- Manage, coach, and grow a team of research scientists and engineers working on biological safety evaluations and classifiers, including hiring, onboarding, performance management, and career development.
- Set the technical direction and roadmap for the biological safety research agenda, including prioritization and decisions about when safeguards are ready to ship.
- Own the quality of capability evaluations assessing what new models can do in the biological domain and turn results into deployment recommendations.
- Guide the development of training and evaluation datasets for safety classifiers with internal and external threat-modeling experts.
- Oversee the training and iteration of safety classifiers alongside ML engineers, optimizing for adversarial robustness and low false-positive rates.
- Ensure investment in tooling and pipelines that make evaluation and classifier development fast and repeatable.
- Establish methods for measuring classifier and evaluation performance against production traffic, identifying gaps, and prioritizing improvements.
- Direct red-teaming and stress-testing of safeguards as threats, models, and product surfaces evolve.
- Partner with Research, Product, Policy, and government affairs teams to embed biological safety throughout the model development lifecycle and serve as an escalation point for biological content.
- Represent the team's work in model cards, blog posts, policy documents, and other external communications.
- Track developments in biology, machine learning, and biosecurity for potential risks and mitigations.
Requirements
- Experience managing a technical team, including hiring, coaching, and performance management.
- A record of setting technical direction and making prioritization decisions under uncertainty.
- Proficiency in Python, with a background in scientific programming and data analysis.
- Solid understanding of machine learning fundamentals sufficient to critically review evaluation design and classifier development.
- Knowledge of modern biology across measurement and engineering, including high-throughput assays, functional characterization, gene synthesis, genome editing, strain construction, and protein engineering.
- Experience designing quantitative experiments or evaluations and drawing defensible conclusions from noisy results.
- Clear analytical and writing skills, with the ability to explain technical concepts to non-technical stakeholders.
- Familiarity with dual-use research concerns and biosecurity frameworks, including select agent regulations, the Biological Weapons Convention, or Australia Group guidelines.
- Comfort with ambiguity and shifting priorities as AI capabilities change.
- Motivation to prevent misuse without obstructing beneficial life sciences work.
Preferred Qualifications
- Three or more years of people management experience, ideally leading research scientists, research engineers, or ML engineers.
- Experience building a team or function from a small headcount, including defining scope, hiring initial team members, and establishing team processes.
- At least eight years of hands-on life sciences experience, with deep expertise in areas such as molecular biology, drug discovery, or computational biology.
- Experience working with large language models, including prompting, fine-tuning, or evaluation.
- Experience training or deploying classifiers or other ML systems in production, including reasoning about precision and recall for rare, high-consequence categories with very low base rates.
- Experience developing ML methods for biological systems or biological data.
- Familiarity with adversarial robustness, red-teaming, or safety evaluation of ML systems.
- Experience leading complex technical projects across multiple stakeholder groups.
Education and Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- Minimum years of experience: Requirements correlate with the internal job level.
- Anthropic currently expects staff to be in one of its offices at least 25% of the time, although some roles may require more office time.
- Anthropic explicitly sponsors visas for eligible roles and candidates and makes reasonable efforts to obtain visas with support from an immigration lawyer.
Compensation
- Annual salary: $405,000–$485,000 USD.
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.