Research Engineer / Performance Engineer, RL Distributed Systems
at Anthropic
USD 500,000-850,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Debugging @ 4
Distributed Systems @ 4
Go @ 7
Kubernetes @ 4
LLM
Machine Learning @ 4
Networking @ 4
Observability @ 4
Python @ 7
Reinforcement Learning @ 3
Rust @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Research Engineer to work on distributed systems that support reinforcement learning at frontier scale. The role involves building systems for training, sampling, and environment execution across large fleets of accelerators and hosts, with a focus on performance, reliability, fault tolerance, and observability.
Responsibilities
- Design, build, and operate distributed systems for reinforcement learning at scale.
- Identify and remove bottlenecks involving scheduling, data movement, storage, networking, and coordination.
- Build fault tolerance through failure detection, isolation, and recovery.
- Design resource management and autoscaling systems that respond to changing demand.
- Build observability systems to diagnose throughput drops, stalls, and unexpected results.
- Create automation for detecting and remediating common problems.
- Design safe operational interfaces for engineers and automated tools.
- Collaborate with researchers and performance engineers to preserve training correctness and avoid subtle nondeterminism.
- Conduct incident reviews, testing, and system redesigns to eliminate classes of failure.
- Write clear design documents and incident writeups.
Requirements
- Strong software engineering skills in Python and at least one systems language, such as Rust, C++, or Go.
- Experience designing, building, and operating large-scale distributed systems in production.
- Deep understanding of distributed systems fundamentals, including consistency, coordination, consensus, failure modes, and recovery.
- Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and networking.
- Experience debugging complex failures across many hosts and services, including failures that cannot be reproduced locally.
- Strong written communication skills.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.
Preferred Qualifications
- Experience running machine learning training or inference infrastructure at scale.
- Experience across scheduling, storage, networking, and orchestration.
- Experience building schedulers, autoscalers, or resource management systems.
- Experience with Kubernetes and sandboxed or virtualized code execution at scale.
- Experience with high-performance networking, RDMA, or collective communication libraries.
- Experience building observability or automated remediation for large fleets.
- Experience with asynchronous Python frameworks such as Trio or asyncio.
- Familiarity with reinforcement learning or large language model training workloads.
Representative Projects
- Designing schedulers for training, sampling, and environment workloads across heterogeneous clusters.
- Building failure detection and recovery systems for long-running jobs.
- Scaling environment execution without increasing training-step tail latency.
- Designing autoscaling policies that rebalance compute as bottlenecks shift.
- Building diagnostics systems that explain throughput reductions and propose fixes.
- Investigating distributed data corruption bugs and redesigning recovery paths.
- Designing operational interfaces that enable automated diagnosis and adjustment under human oversight.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Staff are currently expected to work from one of the company’s offices at least 25% of the time.
More jobs at Anthropic
Staff / Senior Software Engineer, Security Fusion Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Applied AI Engineer, Beneficial Deployments (Life Sciences)
Anthropic · New York City, United States, San Francisco, United States
USD 280,000-320,000 per year
Research Engineer / Research Scientist, RL Frontiers
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 500,000-850,000 per year
Business Systems Analyst, GTM Systems
Anthropic · New York City, United States, San Francisco, United States
USD 270,000-315,000 per year
Applied AI Architect, Partnerships
Anthropic · London, United Kingdom
GBP 150,000-190,000 per year
Similar jobs
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Staff Backend Engineer - Mimir Query, Databases | USA | Remote
Grafana Labs · United States
USD 175,000-210,000 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Germany
PLN 292,500-650,000 per year
Staff Backend Engineer - Mimir Query, Databases
Grafana Labs · Canada
CAD 186,400-223,600 per year
Staff Software Engineer, Infrastructure (Distributed Systems)
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year