Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Compliance
Data Engineering @ 7
Docker @ 3
Kubernetes @ 3
LLM @ 7
Python @ 6
Reinforcement Learning @ 4
Security @ 3
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a senior software engineer to join its RL Data team, which builds systems for producing high-quality reinforcement learning data for Claude. The work includes data collection pipelines, human feedback tooling, execution environments for reinforcement learning tasks, and quality assurance systems that keep training data trustworthy at scale.
This is a foundational role involving architecture decisions, hands-on infrastructure and pipeline engineering, prompt and evaluation iteration, user support, vendor coordination, and end-to-end ownership of systems used by research teams.
Responsibilities
- Own significant parts of the technology stack end-to-end, from technical architecture through operational implementation.
- Build data collection pipelines, review the transcripts they produce, and iterate on prompts, evaluations, and graders.
- Develop and improve quality assurance frameworks to detect reward hacking and ensure execution environment quality.
- Build interfaces that make human data collection efficient and easy for contributors.
- Harden execution environments through sandboxing, snapshotting, and improved tool coverage so tasks operate reliably at training scale.
- Work closely with research teams and domain experts who use the systems.
- Collaborate with operations, security, and compliance partners to roll systems out to new users and vendors.
Requirements
- Track record of owning major projects end-to-end in fast-paced, ambiguous environments, such as as a founder or CTO, forward-deployed engineer, tech lead, founding engineer, or creator of a substantial open-source project.
- Ability to lead and inspire others, plan workstreams, collaborate with cross-functional stakeholders, and proactively eliminate or escalate blockers.
- Strong software engineering skills in at least one modern programming language. Anthropic primarily uses Python and TypeScript and values the ability to learn new tools quickly.
- Effective use of AI tools in day-to-day work.
- Interest in and care for the societal impacts of the work.
- Minimum education of a bachelor's degree or an equivalent combination of education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Preferred Qualifications
- Experience with reinforcement learning for large language models, particularly data-related work involving evaluations, environments, rewards, graders, or training data.
- Experience helping organizations use AI effectively, including integrating third-party tools through APIs, command-line interfaces, and MCP servers.
- Strong data engineering skills, including reliable production pipelines handling large volumes, LLM-powered enrichment, and data quality improvement.
- Experience shipping user-facing products or internal platforms, interviewing users, identifying friction, and measurably improving user experience.
- Basic familiarity with AI safety or security research.
- Familiarity with Docker, Kubernetes, and common cloud infrastructure is a plus.
Representative Projects
- Take a data collection pipeline from research prototype to a production service supporting multiple research teams, including collection, human validation, and grading.
- Develop and harden sandboxed execution environments for long-horizon, high-tool-use agentic tasks across millions of rollouts in a frontier training run.
- Bring new data sources into production training runs while coordinating with product, security, privacy, legal, and infrastructure teams.
- Own the quality assurance layer that determines which tasks enter Claude's training, using automated checks and expert review workflows.
- Reduce the time required to move a task from an initial concept to a production training run by automating, redesigning, or removing bottlenecks.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
Logistics
Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas for this role and states that it will make every reasonable effort to obtain a visa for successful candidates, with support from an immigration lawyer.