Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 3
Debugging
Distributed Systems @ 3
JAX @ 2
Machine Learning @ 2
Profiling
PyTorch @ 2
Reinforcement Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The RL Velocity team owns the efficiency and reliability of Anthropic's RL Science stack, including the infrastructure, tooling, and systems that enable researchers to iterate quickly on training runs. The Research Engineer will build and improve the core platform supporting reinforcement learning at Anthropic, remove bottlenecks, and help the broader organization ship better models faster.
Responsibilities
- Build and improve the RL training infrastructure that researchers use day to day.
- Identify and remove bottlenecks across the RL stack through debugging, profiling, and rearchitecting where needed.
- Partner with researchers and adjacent engineering teams, including inference and sandboxing, to understand pain points and build productivity-enhancing tooling.
- Own the reliability and performance of research runs end to end.
- Contribute to design decisions shaping how Anthropic conducts RL at scale.
Requirements
- Strong software engineering fundamentals and a track record of building performant, reliable systems.
- Experience with ML infrastructure, distributed systems, or research tooling.
- Interest in enabling other people's work and creating leverage through platforms rather than individual experiments.
- Comfort operating across the stack, from low-level performance work to RL algorithms.
- A bias toward shipping and iterating quickly, with a mix of high agency and low ego.
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Experience with large-scale distributed training, including RL, pre-training, or post-training.
- Familiarity with JAX, PyTorch, or similar machine learning frameworks.
- A track record of operating at the edge of research and infrastructure in a fast-moving environment.
Additional Information
Applications are reviewed on a rolling basis, with no application deadline. Years of experience required will correlate with the internal job level requirements for the position. Anthropic expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas, although sponsorship cannot be guaranteed for every role and candidate. The company provides immigration lawyer support and will make every reasonable effort to obtain a visa if an offer is made.