Research Engineer, RL Engineering

USD 500,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI Algorithms @ 3 Distributed Systems @ 3 LLM Machine Learning @ 3 Python @ 3 Reinforcement Learning @ 3

Details

Anthropic is seeking a Research Engineer to build cutting-edge systems for training AI models such as Claude. The role focuses on implementing and improving advanced machine learning techniques and developing the algorithms and infrastructure used by researchers to train reliable, capable, and steerable AI systems.

The Reinforcement Learning Engineering team supports the training of production Claude models and internal research models using RLHF and related methods. This role will build, maintain, and improve the systems and algorithms used for model training, with a focus on performance, reliability, and usability.

Responsibilities

  • Build, maintain, and improve algorithms and systems used for reinforcement learning and model fine-tuning.
  • Improve the speed, reliability, and ease of use of training systems.
  • Profile reinforcement learning pipelines to identify performance improvements.
  • Build systems that regularly launch training jobs in test environments to detect training pipeline problems.
  • Adapt fine-tuning systems to support new model architectures.
  • Build instrumentation to detect and eliminate Python GIL contention in training code.
  • Diagnose and fix training slowdowns occurring after multiple training steps.
  • Implement stable and efficient versions of new training algorithms proposed by researchers.
  • Collaborate closely with researchers and engineers, including through pair programming.

Requirements

  • 4+ years of software engineering experience.
  • A bachelor's degree or equivalent combination of education, training, and experience.
  • Education, training, or professional experience in a field relevant to the role.
  • Interest in building systems and tools that improve the productivity of other people.
  • Results-oriented approach with flexibility and a focus on impact.
  • Interest in learning more about machine learning research.
  • Awareness of the societal impacts of technical work.

Strong candidates may also have experience with:

  • High-performance, large-scale distributed systems.
  • Large-scale large language model training.
  • Python.
  • Implementing large language model fine-tuning algorithms, such as reinforcement learning from human feedback (RLHF).

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration.

Work Arrangement

Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time. Applications are reviewed on a rolling basis.

More jobs at Anthropic

Similar jobs