Research Engineer, RL Scaling Science

GBP 375,000-640,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Debugging @ 3 Machine Learning @ 6 Python @ 5 Reinforcement Learning @ 6

Details

Anthropic’s RL Scaling Science team studies how reinforcement learning behaves as it scales across model size, compute, and task horizon, and turns those findings into training recipes for frontier models. The role combines research and engineering, involving large-scale experiments, long-horizon benchmarks, and validated findings for production training.

Responsibilities

  • Design, run, and interpret large-scale reinforcement learning experiments.
  • Investigate how reinforcement learning improves as horizon, compute, and model size grow.
  • Build and maintain benchmarks for long-horizon reinforcement learning so progress is measurable and reproducible.
  • Translate validated findings into production training recipes and assess when results are robust enough to ship.
  • Debug complex issues at the intersection of research and infrastructure, including failures that appear only at scale.
  • Partner with adjacent reinforcement learning teams across research and engineering to advance the overall RL stack.

Requirements

  • Strong empirical research skills in reinforcement learning, large-scale machine learning training, or a closely adjacent area.
  • Demonstrated ability to own large experiments end-to-end, from design through interpretation.
  • Proficiency in Python and experience working with large-scale or distributed machine learning systems.
  • Comfort operating at the research/systems boundary, including debugging issues where the two meet.
  • Interest in the societal impacts of AI and responsible scaling.
  • Bachelor’s degree or an equivalent combination of education, training, and experience.
  • A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Published or shipped work in long-horizon reinforcement learning or reinforcement learning fundamentals.
  • Experience translating research findings into production training recipes.
  • Demonstrated large-scale industry impact through reinforcement learning interventions.
  • Experience working on frontier-scale training runs with long trajectories.

Representative Projects

  • Design a benchmark suite for long-horizon reinforcement learning that distinguishes genuine capability gains from evaluation artifacts.
  • Stress-test a promising experimental finding across model scales and work with training teams to integrate it into a production recipe.
  • Investigate an unexpected scaling trend in a reinforcement learning run and trace it to a root cause spanning algorithm, data, and infrastructure.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.

Logistics

  • Staff are currently expected to work from one of Anthropic’s offices at least 25% of the time, with some roles requiring more office time.
  • Anthropic sponsors visas and will make reasonable efforts to obtain a visa for candidates who receive an offer, with support from an immigration lawyer.

More jobs at Anthropic

Similar jobs