Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 4
Agentic Systems @ 3
Communication @ 7
Design Patterns
Distributed Systems @ 4
LLM
Machine Learning @ 4
Python @ 6
Reinforcement Learning @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Code RL at Anthropic drives reinforcement learning efforts behind Claude's coding capabilities, creating and scaling agentic coding environments. This engineering role offers significant latitude to set technical direction and standards.
The role involves embedding with research teams, understanding their systems and needs, designing frameworks, APIs, and infrastructure that help researchers move faster, and transferring ownership so teams can maintain those systems. The remit also includes maintaining the health of production reinforcement learning runs through maintainable systems, monitoring, and straightforward triage.
The team's work spans sandboxed execution clients for agentic reinforcement learning environments, large-scale data processing jobs, production dataset lifecycle management, and research environment frameworks.
Responsibilities
- Design widely used APIs, frameworks, and abstractions for engineers and researchers, with attention to interface legibility and principled defaults.
- Embed with research teams on a rotational basis, understand their engineering needs, build supporting systems and APIs, and transfer ownership to the teams.
- Work directly in research codebases to improve reliability and structure without slowing research.
- Anticipate silent failure modes and prevent them through type safety, well-designed invariants, targeted testing, and refactoring.
- Contribute to the reliability of production reinforcement learning systems, including monitoring, regression detection, and triage tooling.
- Help define engineering standards, review practices, and design patterns for a new team.
- Mentor researchers and engineers in adopting engineering standards and patterns.
Requirements
- Deep expertise in Python, including static typing, safe asynchronous and concurrency patterns, and performant Python code.
- A track record of designing intuitive, safe APIs or frameworks adopted by other engineers or teams.
- Experience working productively in large, evolving, or research-oriented codebases that you did not originally write.
- Ability to anticipate failure modes, especially silent failures, and prevent them through system design, type safety, and testing.
- Strong written and verbal communication skills.
- Comfort with ambiguity and the ability to scope work from loosely defined problems through maintainable outcomes.
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Experience building infrastructure, tooling, or frameworks for machine learning research or reinforcement learning workflows.
- Familiarity with reinforcement learning concepts, agentic systems, or large language model training pipelines.
- Experience building or operating large-scale distributed systems.
- Experience building client libraries or SDKs for sandboxed, containerized, or remote execution platforms.
- Experience with large-scale data processing or dataset lifecycle management.
- Experience designing plugin systems or extensible class hierarchies used across an organization.
- Experience embedding with or consulting for other teams and handing off systems for others to own.
- Experience defining code standards, lint rules, or static verification approaches adopted across multiple teams.
- Prior experience as a technical lead or setting engineering standards for a team.
- Prior experience maintaining an open source project.
Representative Projects
- Design a base reinforcement learning environment abstraction that can be subclassed across a wide range of environments.
- Design a model-tool interface for sandboxed agentic environments with explicit serialization semantics.
- Partner with platform teams responsible for the sandbox runtime to specify low-level features that improve the integrity of agentic coding tasks.
- Design probes that detect sandbox regressions early.
- Design the lifecycle and maintenance scheme for a production dataset.
- Lead a research code refactor replacing loosely structured data containers with equivalents that provide stronger correctness guarantees without breaking dependent experiments.
- Design lint rules and code-style requirements favoring statically verifiable patterns and reducing the surface area for silent bugs.
- Build an access layer that enables researchers to discover and reuse data artifacts across teams.
Compensation
- Annual salary: $405,000–$625,000 USD
Logistics and Benefits
- Location-based hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and will make every reasonable effort to obtain a visa for candidates who receive an offer, with assistance from an immigration lawyer.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.