Research Engineer / Research Scientist, RL Frontiers

USD 500,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Algorithms @ 3 Communication @ 3 Debugging @ 3 GPU @ 2 JAX @ 6 Machine Learning @ 3 PyTorch @ 6 Python @ 6 Reinforcement Learning @ 3 Rust @ 3

Details

Anthropic's RL Scaling team studies how reinforcement learning scales as models become larger, episodes become longer, and compute grows. The team develops algorithms and systems that improve throughput, stability, and learning efficiency at frontier scale.

This role spans research and engineering, involving the development of next-generation architectures and reinforcement learning algorithms, scaling them from small-scale experiments to frontier-scale runs, and diagnosing differences across scales.

Responsibilities

  • Study how reinforcement learning training and sampling scale with model size, context length, and compute, and identify algorithmic and systems changes that maintain efficiency.
  • Develop next-generation model architectures and reinforcement learning algorithms and run them efficiently at frontier scale.
  • Scale promising small-scale results to frontier-scale runs and diagnose numerical, algorithmic, or systemic causes of different behavior.
  • Build experimental infrastructure for fast, reproducible comparisons of architecture and algorithm variants at meaningful scale.
  • Own end-to-end performance of the largest reinforcement learning runs, from research code through hardware.
  • Build performance and cost models for proposed architecture and algorithm changes and use them to determine which ideas should be scaled.
  • Investigate training dynamics at scale, including instabilities, divergence, and throughput regressions, and trace them to their root causes.

Requirements

  • Deep familiarity with modern transformer language models, including architecture, training dynamics, and large-scale optimization.
  • Hands-on experience training large models in distributed settings, including data, tensor, and pipeline parallelism tradeoffs.
  • A track record of original technical work in machine learning training or systems, such as new methods, architectures, or optimizations, demonstrated through research, open source, or production impact.
  • Ability to design rigorous experiments at scale, including baselines and ablations, with sufficient statistical care to trust results involving significant compute.
  • Ability to reason quantitatively about the compute, memory, and communication costs of models and algorithms.
  • Strong programming skills in Python and JAX or PyTorch, with the ability to read and modify code at every layer of the stack.

Preferred qualifications

  • Research experience in reinforcement learning, optimization, or large-scale training.
  • Experience developing reinforcement learning algorithms for language models.
  • Experience with scaling laws or other quantitative models of training efficiency.
  • Experience designing or modifying transformer architectures beyond standard configurations.
  • Experience scaling training to large fleets of accelerators and debugging problems that appear only at scale.
  • Deep understanding of numerics in large-scale training, including low-precision formats and sources of instability.
  • Familiarity with how GPU or TPU performance characteristics shape architecture and algorithm choices.
  • Experience with C++ or Rust.

Representative Projects

  • Characterize how a new reinforcement learning algorithm's throughput and learning efficiency change from small models to frontier scale and fix what breaks.
  • Develop a new attention variant, make it work at full scale, and compare its quality and throughput with the baseline.
  • Prepare a large-scale reinforcement learning run by identifying and fixing issues that arise when model size, context length, and compute increase together.
  • Trace a loss instability that appears only beyond a certain scale to its root cause and determine whether the fix belongs in the algorithm, numerics, or system.
  • Build a model that predicts the throughput and cost of a proposed architecture change before kernel implementation.

Education and Logistics

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
  • Required field of study: A field relevant to the role, as demonstrated through coursework, training, or professional experience.
  • Minimum years of experience: Experience requirements correlate with the internal job level.
  • Anthropic expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
  • Anthropic sponsors visas, although sponsorship cannot be successfully provided for every role and candidate.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration.

More jobs at Anthropic

Similar jobs