Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Debugging @ 3
Machine Learning @ 6
Python @ 5
Reinforcement Learning @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s RL Scaling Science team studies how reinforcement learning behaves as it scales across model size, compute, and task horizon, and turns those findings into training recipes for frontier models. The role combines research and engineering, involving large-scale experiments, long-horizon benchmarks, and validated findings for production training.
Responsibilities
- Design, run, and interpret large-scale reinforcement learning experiments.
- Investigate how reinforcement learning improves as horizon, compute, and model size grow.
- Build and maintain benchmarks for long-horizon reinforcement learning so progress is measurable and reproducible.
- Translate validated findings into production training recipes and assess when results are robust enough to ship.
- Debug complex issues at the intersection of research and infrastructure, including failures that appear only at scale.
- Partner with adjacent reinforcement learning teams across research and engineering to advance the overall RL stack.
Requirements
- Strong empirical research skills in reinforcement learning, large-scale machine learning training, or a closely adjacent area.
- Demonstrated ability to own large experiments end-to-end, from design through interpretation.
- Proficiency in Python and experience working with large-scale or distributed machine learning systems.
- Comfort operating at the research/systems boundary, including debugging issues where the two meet.
- Interest in the societal impacts of AI and responsible scaling.
- Bachelor’s degree or an equivalent combination of education, training, and experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Published or shipped work in long-horizon reinforcement learning or reinforcement learning fundamentals.
- Experience translating research findings into production training recipes.
- Demonstrated large-scale industry impact through reinforcement learning interventions.
- Experience working on frontier-scale training runs with long trajectories.
Representative Projects
- Design a benchmark suite for long-horizon reinforcement learning that distinguishes genuine capability gains from evaluation artifacts.
- Stress-test a promising experimental finding across model scales and work with training teams to integrate it into a production recipe.
- Investigate an unexpected scaling trend in a reinforcement learning run and trace it to a root cause spanning algorithm, data, and infrastructure.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.
Logistics
- Staff are currently expected to work from one of Anthropic’s offices at least 25% of the time, with some roles requiring more office time.
- Anthropic sponsors visas and will make reasonable efforts to obtain a visa for candidates who receive an offer, with support from an immigration lawyer.
More jobs at Anthropic
Software Engineer, Business Technology
Anthropic · London, United Kingdom
GBP 255,000-325,000 per year
Repairs Program Lead - Data Center Operations
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Product Policy Manager, Product Risk
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 245,000-285,000 per year
Staff + Senior Software Engineer, Scaling
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Applied AI Engineer, Beneficial Deployments (Life Sciences)
Anthropic · New York City, United States, San Francisco, United States
USD 280,000-320,000 per year
Similar jobs
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Senior Deep Learning Algorithm Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Applied Deep Learning PhD Research Intern, Reinforcement Learning for LLMs - Fall 2026
Nvidia · Santa Clara, United States
USD 30-94 per hour
Senior Robotics Research Engineer, Robotics and AI for Drug Discovery
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff - Imagine Model
SpaceXAI · Palo Alto, United States, Seattle, United States
USD 180,000-440,000 per year
Research Scientist, Robotics Research - PhD New College Grad 2026
Nvidia · Seattle, United States
USD 168,000-264,500 per year