Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 6
Debugging
Distributed Systems @ 4
JAX @ 3
Machine Learning @ 3
Profiling
PyTorch @ 3
Reinforcement Learning
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society.
The RL Velocity team owns the efficiency and reliability of Anthropic's RL Science stack, including the infrastructure, tooling, and systems that enable researchers to iterate quickly on training runs. This role focuses on building and improving the core platform underpinning reinforcement learning at Anthropic, removing research bottlenecks, and helping the organization ship better models faster.
Responsibilities
- Build and improve the RL training infrastructure used by researchers day to day.
- Identify and remove bottlenecks across the RL stack through debugging, profiling, and rearchitecting.
- Partner with researchers and adjacent engineering teams, including inference and sandboxing, to understand pain points and deliver productivity-enhancing tooling.
- Own the reliability and performance of research runs end to end.
- Contribute to design decisions shaping how Anthropic performs RL at scale.
Requirements
- Strong software engineering fundamentals and a track record of building performant, reliable systems.
- Experience with ML infrastructure, distributed systems, or research tooling.
- Interest in enabling other people's work and creating leverage through platforms rather than individual experiments.
- Comfort operating across the stack, from low-level performance work to RL algorithms.
- A bias toward shipping and iterating quickly, with a mix of high agency and low ego.
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Relevant education, training, or professional experience in a field related to the role.
Preferred Qualifications
- Experience with large-scale distributed training, including RL, pre-training, or post-training.
- Familiarity with JAX, PyTorch, or similar machine learning frameworks.
- A track record of working at the intersection of research and infrastructure in a fast-moving environment.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from one of the company's offices at least 25% of the time, although some roles may require more office time.
Anthropic explicitly sponsors visas for eligible roles and candidates and makes reasonable efforts to support visa applications with the assistance of an immigration lawyer.