Anthropic Fellows Program, Reinforcement Learning

USD 200,200 per year
MIDDLE
✅ Remote ✅ On-site

Tech Stack

AI @ 3 API Algorithms Distributed Systems @ 3 HPC LLM Machine Learning @ 6 Mathematics @ 6 Python @ 5 Reinforcement Learning Security

Details

Anthropic’s Fellows Program fosters AI research and engineering talent by providing funding and mentorship to promising technical talent, regardless of previous experience. Fellows primarily use external infrastructure, such as open-source models and public APIs, to work on empirical projects aligned with Anthropic’s research priorities, with the goal of producing a public output such as a paper submission.

The program is full-time for four months, with the next cohort expected to start November 2. Fellows receive direct mentorship from Anthropic researchers, access to shared workspaces in Berkeley or London, connection to the broader AI safety and security research community, funding for compute of approximately $15,000 per month, and funding for other research expenses.

Responsibilities

  • Build model-based tools to better understand AI training data and improve training data quality.
  • Conduct research into generalization.
  • Create reinforcement learning environments to improve Claude models at capabilities within the fellow’s domain of expertise.
  • Build reinforcement learning environments for safety-related tasks.
  • Research and implement solutions involving reinforcement learning algorithms.
  • Implement ideas quickly and communicate clearly.
  • Analyze and debug model training processes.
  • Balance research exploration with engineering rigor and operational reliability.

Requirements

  • Fluency in Python programming.
  • Availability to work full-time for the duration of the Fellows Program.
  • A strong technical background in computer science, mathematics, physics, or a related discipline.
  • Strong software engineering skills and experience building complex machine learning systems.
  • Ability to collaborate across research and engineering disciplines.
  • Comfort working with large-scale distributed systems and high-performance computing.
  • Experience training, fine-tuning, or evaluating large language models.
  • Motivation to help ensure that AI is safe and beneficial for society.
  • Work authorization in the United States, United Kingdom, or Canada, with the ability to be located in that country during the program.

Logistics

Designated shared workspaces are available in London and Berkeley, and remote fellows are also accepted in the United Kingdom, United States, or Canada. Fellows must have or independently obtain full-time work authorization in one of these countries. Anthropic is not currently able to sponsor visas for fellows.

The program runs for four months on a full-time basis, with possible extensions. Anthropic does not guarantee full-time employment offers after the program.

Compensation

The expected base stipend is 3,850 USD, 2,310 GBP, or 4,300 CAD per week, with an expectation of 40 hours per week. Benefits vary by country.

More jobs at Anthropic

Similar jobs