Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Communication @ 7
Fraud @ 4
Machine Learning
Prompt Engineering @ 4
Python @ 6
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking software engineers to help build safety and oversight mechanisms for its AI systems. As a software engineer on the Safeguards team, you will work to monitor models, prevent misuse, and ensure user well-being. The role focuses on building systems to detect unwanted model behaviors and prevent disallowed use of models while enforcing terms of service and acceptable use policies.
Responsibilities
- Develop monitoring systems to detect unwanted behaviors from API partners and potentially take automated enforcement actions.
- Surface detected behaviors in internal dashboards for analyst manual review.
- Build abuse detection mechanisms and infrastructure.
- Surface abuse patterns to research teams to help harden models during training.
- Build robust, reliable, multi-layered defenses for real-time improvement of safety mechanisms at scale.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or comparable experience.
- Proficiency in Python and TypeScript.
- Ability to work across the stack.
- Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.
Preferred Qualifications
- 8+ years of experience in a software engineering position.
- Experience with integrity, spam, fraud, or abuse detection and mitigation.
- Experience building trust and safety detection mechanisms and interventions for AI/ML systems.
- Experience with prompt engineering, jailbreak attacks, and other adversarial inputs.
- Experience working closely with operational teams to build custom internal tooling.
Compensation
Annual salary: $320,000–$485,000 USD.
Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- Minimum years of experience correlate with the internal job-level requirements.
- Hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more in-office time.
- Anthropic sponsors visas, although sponsorship is not available for every role and candidate. The company makes reasonable efforts to obtain visas and retains an immigration lawyer to assist with the process.
About Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Anthropic is a public benefit corporation headquartered in San Francisco. The company offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.