Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Data Pipelines
LLM
Machine Learning @ 6
Reinforcement Learning @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Domain Scaling team works to make Claude world-class at real-world knowledge work in domains such as finance, healthcare, and legal. This role combines applied research and data sourcing, including real-world and synthetic data, to improve model capabilities. The role owns the end-to-end process of creating reinforcement learning environments for new capabilities, including identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance.
Responsibilities
- Own the data strategy for knowledge work verticals, from task sourcing through reinforcement learning training.
- Manage technical relationships with external data vendors, including evaluating data quality and designing rewards.
- Collaborate with domain experts to design data pipelines and evaluations.
- Explore novel approaches for creating reinforcement learning environments for high-value tasks.
- Develop and improve quality assurance frameworks to detect reward hacking and ensure environment quality.
- Run generalization experiments to measure how changes in data strategy affect model capabilities.
- Partner with other reinforcement learning research teams and product teams to translate capability goals into training environments and evaluations.
Requirements
- Experience fine-tuning large language models for specific domains or real-world use cases.
- Experience with reinforcement learning, reward design, or training data curation for large language models.
- Ability to manage technical vendor relationships and iterate quickly based on feedback.
- Comfort reading datasets to understand them and identify issues.
- Strong cross-functional collaboration skills.
- Passion for making AI more useful and accessible across different industries.
- Interest in a role combining applied research and hands-on data work.
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role, as demonstrated through coursework, training, or professional experience.
- Minimum years of experience will correlate with the internal job-level requirements for the position.
Strong candidates may also have experience training production machine learning systems, designing evaluations or benchmarks for large language models, domain expertise in a relevant vertical, or working with external vendors or technical partners.
Company And Research Context
Anthropic's mission is to create reliable, interpretable, and steerable AI systems. The company works collaboratively on large-scale AI research efforts, including research related to GPT-3, circuit-based interpretability, multimodal neurons, scaling laws, AI and compute, concrete problems in AI safety, and learning from human preferences.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.