Research Engineer / Performance Engineer, RL Distributed Systems

USD 500,000-850,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Communication @ 7 Debugging @ 4 Distributed Systems @ 4 Go @ 7 Kubernetes @ 4 LLM Machine Learning @ 4 Networking @ 4 Observability @ 4 Python @ 7 Reinforcement Learning @ 3 Rust @ 7

Details

Anthropic is seeking a Research Engineer to work on distributed systems that support reinforcement learning at frontier scale. The role involves building systems for training, sampling, and environment execution across large fleets of accelerators and hosts, with a focus on performance, reliability, fault tolerance, and observability.

Responsibilities

  • Design, build, and operate distributed systems for reinforcement learning at scale.
  • Identify and remove bottlenecks involving scheduling, data movement, storage, networking, and coordination.
  • Build fault tolerance through failure detection, isolation, and recovery.
  • Design resource management and autoscaling systems that respond to changing demand.
  • Build observability systems to diagnose throughput drops, stalls, and unexpected results.
  • Create automation for detecting and remediating common problems.
  • Design safe operational interfaces for engineers and automated tools.
  • Collaborate with researchers and performance engineers to preserve training correctness and avoid subtle nondeterminism.
  • Conduct incident reviews, testing, and system redesigns to eliminate classes of failure.
  • Write clear design documents and incident writeups.

Requirements

  • Strong software engineering skills in Python and at least one systems language, such as Rust, C++, or Go.
  • Experience designing, building, and operating large-scale distributed systems in production.
  • Deep understanding of distributed systems fundamentals, including consistency, coordination, consensus, failure modes, and recovery.
  • Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and networking.
  • Experience debugging complex failures across many hosts and services, including failures that cannot be reproduced locally.
  • Strong written communication skills.
  • Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

Preferred Qualifications

  • Experience running machine learning training or inference infrastructure at scale.
  • Experience across scheduling, storage, networking, and orchestration.
  • Experience building schedulers, autoscalers, or resource management systems.
  • Experience with Kubernetes and sandboxed or virtualized code execution at scale.
  • Experience with high-performance networking, RDMA, or collective communication libraries.
  • Experience building observability or automated remediation for large fleets.
  • Experience with asynchronous Python frameworks such as Trio or asyncio.
  • Familiarity with reinforcement learning or large language model training workloads.

Representative Projects

  • Designing schedulers for training, sampling, and environment workloads across heterogeneous clusters.
  • Building failure detection and recovery systems for long-running jobs.
  • Scaling environment execution without increasing training-step tail latency.
  • Designing autoscaling policies that rebalance compute as bottlenecks shift.
  • Building diagnostics systems that explain throughput reductions and propose fixes.
  • Investigating distributed data corruption bugs and redesigning recovery paths.
  • Designing operational interfaces that enable automated diagnosis and adjustment under human oversight.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Staff are currently expected to work from one of the company’s offices at least 25% of the time.

More jobs at Anthropic

Similar jobs