Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 3
Distributed Systems @ 3
LLM
Machine Learning @ 3
Python @ 3
Reinforcement Learning @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Research Engineer to build cutting-edge systems for training AI models such as Claude. The role focuses on implementing and improving advanced machine learning techniques and developing the algorithms and infrastructure used by researchers to train reliable, capable, and steerable AI systems.
The Reinforcement Learning Engineering team supports the training of production Claude models and internal research models using RLHF and related methods. This role will build, maintain, and improve the systems and algorithms used for model training, with a focus on performance, reliability, and usability.
Responsibilities
- Build, maintain, and improve algorithms and systems used for reinforcement learning and model fine-tuning.
- Improve the speed, reliability, and ease of use of training systems.
- Profile reinforcement learning pipelines to identify performance improvements.
- Build systems that regularly launch training jobs in test environments to detect training pipeline problems.
- Adapt fine-tuning systems to support new model architectures.
- Build instrumentation to detect and eliminate Python GIL contention in training code.
- Diagnose and fix training slowdowns occurring after multiple training steps.
- Implement stable and efficient versions of new training algorithms proposed by researchers.
- Collaborate closely with researchers and engineers, including through pair programming.
Requirements
- 4+ years of software engineering experience.
- A bachelor's degree or equivalent combination of education, training, and experience.
- Education, training, or professional experience in a field relevant to the role.
- Interest in building systems and tools that improve the productivity of other people.
- Results-oriented approach with flexibility and a focus on impact.
- Interest in learning more about machine learning research.
- Awareness of the societal impacts of technical work.
Strong candidates may also have experience with:
- High-performance, large-scale distributed systems.
- Large-scale large language model training.
- Python.
- Implementing large language model fine-tuning algorithms, such as reinforcement learning from human feedback (RLHF).
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration.
Work Arrangement
Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time. Applications are reviewed on a rolling basis.