Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Agentic Systems @ 4
Communication @ 6
Debugging
Distributed Systems @ 4
LLM
Machine Learning @ 4
Reinforcement Learning @ 4
Statistics @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Staff Research Engineer to study how large teams of AI agents scale, including how performance, cost, and coordination change with the number of agents, compute budget, and task length. The role combines research and engineering, involving large-scale experiments, platform development, evaluations, and debugging complex systems.
Responsibilities
- Design, run, and interpret large-scale experiments on agent teams.
- Investigate how performance and efficiency change as team size, compute, and task horizon grow.
- Identify bottlenecks that limit scaling.
- Build and scale reliable systems for running very large agent teams.
- Debug failures that emerge at scale.
- Design evaluations for long-horizon problems and ensure their results remain trustworthy.
- Build tooling and metrics that help researchers understand agent-team behavior.
- Partner with research teams across Anthropic so they can run experiments on the platform.
- Communicate research findings clearly.
Requirements
- Significant software engineering, machine learning, or research engineering experience.
- Experience owning a substantial project end to end, such as a large system, evaluation or benchmark, agent product, or research project.
- Ability and interest to work across both research and engineering.
- Quantitative reasoning about complex systems and careful evaluation of data.
- Ability to work from vague questions rather than a detailed specification.
- Results-oriented approach with flexibility and focus on impact.
- Clear written and verbal communication.
- Care for the societal impacts of the work.
- A bachelor's degree or equivalent combination of education, training, and experience. The field of study should be relevant to the role through coursework, training, or professional experience.
Preferred Qualifications
- Experience building or operating large-scale distributed systems, including schedulers, sandboxed code execution, inference infrastructure, or reinforcement learning infrastructure.
- Experience building evaluations, benchmarks, or harnesses for large language models or agents.
- Experience building complex agentic systems that use large language models.
- Experience with scaling laws or other large-scale empirical research.
- Background in operations research, statistics, economics, physics, quantitative finance, or another field focused on modeling and optimizing complex systems.
Formal certifications, academic research experience, publication history, and prior experience with multi-agent systems or reinforcement learning are not required.
Representative Projects
- Measuring how performance scales with the number of agents on a difficult problem and explaining where and why the scaling curve changes.
- Preparing large agent runs by identifying and fixing failures that emerge as team size and task length increase.
- Allocating a fixed compute budget across a team of agents to solve a problem efficiently.
- Building tooling that makes the activity of a large agent team understandable to researchers.
- Designing evaluations that distinguish genuine teamwork improvements from artifacts of the evaluation setup.
Compensation
- Annual salary: $500,000–$850,000 USD.
Work Location and Benefits
- Hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
- Anthropic sponsors visas and makes reasonable efforts to obtain visas for candidates, with support from an immigration lawyer.
- Anthropic is a public benefit corporation headquartered in San Francisco.
More jobs at Anthropic
Program Manager, Safeguards Policy, Enforcement, and Threat Intelligence
Anthropic · San Francisco, United States
USD 285,000-330,000 per year
Security Audit & Controls, Security GRC
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 270,000-345,000 per year
Security Audit & Controls, Security GRC
Anthropic · New York City, United States, San Francisco, United States
USD 0 per year
Technical Program Manager, Life Sciences
Anthropic · New York City, United States, San Francisco, United States
USD 290,000-365,000 per year
Software Engineer, Sandboxing
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-485,000 per year
Similar jobs
Staff Software Engineer, Environments Infrastructure
Anthropic · New York City, United States, San Francisco, United States
USD 405,000-605,000 per year
Research Engineer, Model Evaluations
Anthropic · New York City, United States, San Francisco, United States
USD 500,000-850,000 per year
Database Research Scientist
ClickHouse · Canada, Germany, United Kingdom, Netherlands, United States
USD 175,000-235,000 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Research Engineer / Performance Engineer, RL Distributed Systems
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 500,000-850,000 per year
Staff Software Engineer, Code RL
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-625,000 per year
Research Engineer, Knowledge Team
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 350,000-850,000 per year