Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CUDA @ 3
GPU @ 3
LLM @ 3
Performance Optimization @ 3
Profiling @ 3
PyTorch @ 3
Python @ 6
Reinforcement Learning @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is hiring a Research Engineer for its Code RL team within the Reinforcement Learning organization. The role focuses on advancing models' ability to write, edit, test, debug, and ship software end to end across real codebases and tools, correctly, quickly, and safely.
The role combines research and engineering. The engineer will design reinforcement learning environments and coding tasks, build reward signals and verifiers that measure code quality, run training experiments on frontier models, diagnose model performance on software-engineering tasks, and improve the speed and reliability of the supporting pipelines. Focus areas include agentic coding behaviors, code correctness, long-horizon autonomous engineering, and high-performance code for accelerators.
Responsibilities
- Develop systems that enable models to use computers effectively.
- Advance code generation through reinforcement learning.
- Design RL environments and coding tasks.
- Build reward signals and verifiers that capture what constitutes good code.
- Run training experiments on frontier models.
- Diagnose why models do or do not improve at software-engineering tasks.
- Improve the speed and reliability of research and training pipelines.
- Contribute to fundamental RL research for large language models.
- Build scalable RL infrastructure and training methodologies.
- Enhance model reasoning capabilities.
- Collaborate with alignment, frontier red-team, and applied production training teams.
Requirements
- Strong software-engineering skills and deep Python expertise, including asynchronous and concurrent programming.
- Ability to own systems end to end and debug across the stack.
- Ability to balance research exploration with engineering implementation.
- Ability to contribute rigorously to experimental design and interpretation of results.
- Commitment to code quality, testing, and performance.
- Passion for the potential impact of AI and commitment to developing safe and beneficial systems.
- Bachelor's degree or an equivalent combination of education, training, and experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
Strong candidates may also have experience with:
- Reinforcement learning, RLHF, post-training, or LLM fine-tuning.
- Coding agents, code-execution sandboxes, evaluation harnesses, verifiers, or developer tooling.
- Program analysis, testing, verification, compilers, or formal methods.
- PyTorch and large-scale distributed training.
- Performance profiling and optimization of machine-learning systems.
- CUDA, GPU, or TPU kernels and accelerator-performance optimization.
- Virtualization and sandboxed code-execution environments.
Benefits
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
- Visa sponsorship may be available; Anthropic states that it sponsors visas and makes reasonable efforts to obtain visas for candidates who receive an offer.
Work Arrangement
The role is subject to a location-based hybrid policy. Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.