Research Engineer, Machine Learning (Reinforcement Learning)

GBP 260,000-630,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 API Automated Testing Communication @ 6 Debugging Distributed Systems @ 3 GPU HPC JAX @ 3 Kubernetes @ 3 LLM @ 2 Machine Learning @ 3 Mathematics Profiling PyTorch @ 3 Python @ 5 Reinforcement Learning @ 3 Rust @ 3 TensorFlow @ 3

Details

Anthropic is seeking a Research Engineer to advance the capabilities and safety of large language models within its Reinforcement Learning teams. The role combines research and engineering, including implementing novel approaches, contributing to research direction, developing agentic models through tool use, improving reasoning capabilities, and building prototypes for internal use, productivity, and evaluation.

Responsibilities

  • Architect and optimize reinforcement learning infrastructure, including training abstractions and distributed experiment management across GPU clusters.
  • Design, implement, and test training environments, evaluations, and methodologies for reinforcement learning agents.
  • Improve performance through profiling, optimization, benchmarking, caching, and debugging distributed systems.
  • Collaborate with research and engineering teams to develop automated testing frameworks, clean APIs, and scalable infrastructure for AI research.
  • Contribute to research involving computer use, autonomous software generation, mathematics, reasoning, and large language models.

Requirements

  • Proficiency in Python and asynchronous or concurrent programming, including frameworks such as Trio.
  • Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Industry experience in machine learning research.
  • Ability to balance research exploration with engineering implementation.
  • Strong systems design, communication, code quality, testing, and performance skills.
  • Passion for developing safe and beneficial AI systems.
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience. Relevant fields of study may be demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Familiarity with LLM architectures and training methodologies.
  • Experience with reinforcement learning techniques and environments.
  • Experience with virtualization and sandboxed code execution environments.
  • Experience with Kubernetes, distributed systems, or high-performance computing.
  • Experience with Rust and/or C++.

Formal certifications, education credentials, academic research experience, and publication history are not required for strong candidates.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.

Work Policy and Sponsorship

The role is based in London, UK, with a hybrid policy requiring staff to work from an Anthropic office at least 25% of the time. Anthropic sponsors visas where possible and makes reasonable efforts to support visa applications with immigration counsel.

More jobs at Anthropic

Similar jobs