Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Deep Learning @ 5
Distributed Systems @ 3
HPC
LLM
Machine Learning @ 6
Python @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's production models undergo sophisticated post-training processes to enhance their capabilities, alignment, and safety. As a Research Engineer on the Post-Training team, you will train base models through the complete post-training stack to deliver production Claude models.
The role combines cutting-edge research with production engineering, including implementing, scaling, and improving Constitutional AI, RLHF, and other alignment methodologies. The work directly impacts the quality, safety, and capabilities of production models.
Interviews are conducted in Python. The role may require responding to incidents on short notice, including weekends.
Responsibilities
- Implement and optimize post-training techniques at scale on frontier models.
- Conduct research to develop and optimize post-training recipes that improve production model quality.
- Design, build, and run robust, efficient pipelines for model fine-tuning and evaluation.
- Develop tools to measure and improve model performance across various dimensions.
- Collaborate with research teams to translate emerging techniques into production-ready implementations.
- Debug complex issues in training pipelines and model behavior.
- Help establish best practices for reliable, reproducible model post-training.
Requirements
- Strong software engineering skills and experience building complex machine learning systems.
- Experience with large-scale distributed systems and high-performance computing.
- Experience training, fine-tuning, or evaluating large language models.
- Ability to balance research exploration with engineering rigor and operational reliability.
- Ability to analyze and debug model training processes.
- Proficiency in Python, deep learning frameworks, and distributed computing.
- Ability to collaborate across research and engineering disciplines and work effectively in ambiguous, fast-moving research environments.
- A bachelor's degree or equivalent combination of education, training, and/or experience.
- Education, training, or professional experience in a field relevant to the role.
Strong candidates may also have experience with LLMs and a keen interest in AI safety and responsible deployment. Candidates are welcomed at various experience levels, with a preference for senior engineers who have hands-on experience with frontier AI systems. Years of experience will correlate with the internal job level requirements.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.