Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 3
Communication @ 3
Debugging @ 3
GPU @ 2
JAX @ 6
Machine Learning @ 3
PyTorch @ 6
Python @ 6
Reinforcement Learning @ 3
Rust @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's RL Scaling team studies how reinforcement learning scales as models become larger, episodes become longer, and compute grows. The team develops algorithms and systems that improve throughput, stability, and learning efficiency at frontier scale.
This role spans research and engineering, involving the development of next-generation architectures and reinforcement learning algorithms, scaling them from small-scale experiments to frontier-scale runs, and diagnosing differences across scales.
Responsibilities
- Study how reinforcement learning training and sampling scale with model size, context length, and compute, and identify algorithmic and systems changes that maintain efficiency.
- Develop next-generation model architectures and reinforcement learning algorithms and run them efficiently at frontier scale.
- Scale promising small-scale results to frontier-scale runs and diagnose numerical, algorithmic, or systemic causes of different behavior.
- Build experimental infrastructure for fast, reproducible comparisons of architecture and algorithm variants at meaningful scale.
- Own end-to-end performance of the largest reinforcement learning runs, from research code through hardware.
- Build performance and cost models for proposed architecture and algorithm changes and use them to determine which ideas should be scaled.
- Investigate training dynamics at scale, including instabilities, divergence, and throughput regressions, and trace them to their root causes.
Requirements
- Deep familiarity with modern transformer language models, including architecture, training dynamics, and large-scale optimization.
- Hands-on experience training large models in distributed settings, including data, tensor, and pipeline parallelism tradeoffs.
- A track record of original technical work in machine learning training or systems, such as new methods, architectures, or optimizations, demonstrated through research, open source, or production impact.
- Ability to design rigorous experiments at scale, including baselines and ablations, with sufficient statistical care to trust results involving significant compute.
- Ability to reason quantitatively about the compute, memory, and communication costs of models and algorithms.
- Strong programming skills in Python and JAX or PyTorch, with the ability to read and modify code at every layer of the stack.
Preferred qualifications
- Research experience in reinforcement learning, optimization, or large-scale training.
- Experience developing reinforcement learning algorithms for language models.
- Experience with scaling laws or other quantitative models of training efficiency.
- Experience designing or modifying transformer architectures beyond standard configurations.
- Experience scaling training to large fleets of accelerators and debugging problems that appear only at scale.
- Deep understanding of numerics in large-scale training, including low-precision formats and sources of instability.
- Familiarity with how GPU or TPU performance characteristics shape architecture and algorithm choices.
- Experience with C++ or Rust.
Representative Projects
- Characterize how a new reinforcement learning algorithm's throughput and learning efficiency change from small models to frontier scale and fix what breaks.
- Develop a new attention variant, make it work at full scale, and compare its quality and throughput with the baseline.
- Prepare a large-scale reinforcement learning run by identifying and fixing issues that arise when model size, context length, and compute increase together.
- Trace a loss instability that appears only beyond a certain scale to its root cause and determine whether the fix belongs in the algorithm, numerics, or system.
- Build a model that predicts the throughput and cost of a proposed architecture change before kernel implementation.
Education and Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role, as demonstrated through coursework, training, or professional experience.
- Minimum years of experience: Experience requirements correlate with the internal job level.
- Anthropic expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas, although sponsorship cannot be successfully provided for every role and candidate.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration.
More jobs at Anthropic
Staff / Senior Software Engineer, Security Fusion Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Applied AI Engineer, Beneficial Deployments (Life Sciences)
Anthropic · New York City, United States, San Francisco, United States
USD 280,000-320,000 per year
Research Engineer / Performance Engineer, RL Distributed Systems
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 500,000-850,000 per year
Business Systems Analyst, GTM Systems
Anthropic · New York City, United States, San Francisco, United States
USD 270,000-315,000 per year
Applied AI Architect, Partnerships
Anthropic · London, United Kingdom
GBP 150,000-190,000 per year
Similar jobs
Member of Technical Staff - Imagine Model
SpaceXAI · Palo Alto, United States, Seattle, United States
USD 180,000-440,000 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Research Intern, Robotics – Summer 2027
Nvidia · Seattle, United States
USD 38-94 per hour
Research Scientist, Networking Research - PhD New College Grad 2026
Nvidia · Santa Clara, United States
USD 168,000-264,500 per year
Software Engineering Manager, Robotics Neural Reconstruction and Real2Sim Applications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Member of Technical Staff - Multimodal Understanding
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Senior Robotics Research Engineer, Robotics and AI for Drug Discovery
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Research Scientist, Robotics Research - PhD New College Grad 2026
Nvidia · Seattle, United States
USD 168,000-264,500 per year