Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 3
AWS @ 3
Audit
CI/CD @ 3
Communication @ 6
Docker @ 3
GCP @ 3
LLM @ 2
Machine Learning
Observability
Python @ 5
React @ 5
Reinforcement Learning
TypeScript @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. The Reinforcement Learning (RL) organization trains Claude to be capable, reliable, and safe.
As a Full-Stack Software Engineer in RL, you will build the platforms, tools, and interfaces that power environment creation, data collection, and training observability. The quality of Claude’s next generation depends on the quality of the data trained on, and the systems you build are what make that data possible.
You will own product surfaces end-to-end—from backend services and APIs to web UIs used by researchers, external vendors, and thousands of data labelers. You don’t need a background in ML research; what matters is shipping a polished, reliable product quickly while dealing with ambiguity in a high-stakes environment.
Responsibilities
- Build and extend web platforms for RL environment creation, management, and quality review, including environment configuration, versioning, and validation workflows
- Develop vendor-facing interfaces and tooling enabling external partners to create, submit, and iterate on training environments with minimal friction
- Design and implement platforms for human data collection at scale, including labeling workflows, quality assurance systems, and feedback mechanisms that surface reward signal integrity issues early
- Build evaluation dashboards and observability UIs providing researchers real-time insight into environment quality, training run health, and reward hacking
- Create backend services and APIs connecting environment authoring tools, data collection systems, and RL training infrastructure
- Build and expand scalable code data generation pipelines producing diverse programming tasks with robust reward signals across languages and difficulty levels
- Develop onboarding automation and documentation tooling so new vendors and internal users ramp up in hours, not weeks
- Partner closely with RL researchers, data operations, and vendor management to translate ambiguous requirements into well-scoped, well-designed products
Requirements
- Strong software engineering fundamentals and real full-stack range (owning a surface from database schema to frontend)
- Proficient in Python and a modern web stack (React, TypeScript, or similar)
- Track record of shipping systems that solved a hard problem (e.g., built something that made a team significantly faster)
- High agency: identify what needs to be done and drive it forward without waiting for a ticket
- Strong communication skills; can turn vague asks into well-scoped work and collaborate with researchers, operations teams, and engineers
- Ability to build interfaces intuitive for both technical researchers and non-technical labelers
- Thrive in a fast-moving environment where priorities shift
- Care about Anthropic’s mission to build safe, beneficial AI and want your work to contribute directly
Strong Candidates May Also Have
- Experience building data collection, labeling, or annotation platforms, ideally at scale across many vendors or task types
- Background building multi-tenant platforms with role-based access, audit trails, and vendor management workflows
- Experience with cloud infrastructure (GCP or AWS), Docker, and CI/CD pipelines
- Familiarity with LLM training, fine-tuning, or evaluation workflows
- Experience with async Python (Trio, asyncio) or high-throughput API design
- Background in dashboards, monitoring, or observability tooling
- Experience working directly with external vendors/partners on technical integrations
- Non-linear backgrounds (e.g., math/physics into SWE, competitive programming, research into engineering, or side projects that outgrew scope)
Logistics
- Location-based hybrid policy: expect all staff to be in one of the offices at least 25% of the time
- Visa sponsorship: Anthropic does sponsor visas and will make every reasonable effort to get candidates a visa if they make an offer (and retains an immigration lawyer)