Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
A/B Testing @ 3
LLM @ 3
Python @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Product Designer, Evals & Prompts to join the Product Prompt and Eval Design team. The team designs prompts and evaluations for Claude's product surfaces, ensuring that the model and product remain aligned with user expectations, product strategy, and safety requirements across model launches.
This role focuses on building evaluations, the harness that runs them, and tools that enable designers to perform evaluation work independently. The role works with product surface owners, engineers, and prompt engineers during model releases.
Responsibilities
- Write and revise prompts behind Claude's tools, features, and behaviors on product surfaces.
- Test product surfaces, turn findings into prompt fixes, ship them, and confirm that users receive the intended prompts.
- Build graders that validate prompt fixes and rerun evaluations on subsequent models.
- Convert designers' manual rubrics into automated evaluations and review transcripts to identify gaps in evaluation coverage.
- Build visual, low-code evaluation tools for non-engineering users, including tools to assemble comparison sets from real transcripts, convert plain-English rubrics into graders, compare prompt variants across models, and review results without a notebook.
- Observe designers using evaluation tools and simplify the workflows.
- Support model releases by testing each surface against new models, writing prompt fixes and migrations, and creating prompts for features launching with each model.
- Establish and scale the evaluation harness, including a test environment for 50 to 100 tools with trustworthy settings.
- Maintain evaluations across models and determine whether regressions originate in the harness or the model.
- Package behaviors that prompting cannot fix for training, including graders, human-feedback questions, and good/bad preference pairs for qualities such as writing quality.
Requirements
- Production-quality Python experience.
- Experience building and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, regression suites, and the infrastructure to run evaluations across models.
- Experience building internal tools with interfaces for people who do not write code.
- Experience establishing test harnesses, sandboxing tool calls, and pinning settings to ensure comparable runs.
- Experience shipping prompts or working closely with prompt developers, with an understanding of why a prompt that works on one model may fail on another.
- Ability to read transcripts in addition to reviewing scores.
Preferred Qualifications
- Experience working within a model-launch cycle.
- A/B testing experience and the ability to connect offline evaluations to online outcomes.
- Front-end or notebook-to-application experience, with an understanding of how to make evaluation results immediately legible.
- Experience turning product rubrics into training signals, including graders, human-feedback questions, or preference pairs.
- Care about how Claude behaves for users, not only whether a metric changes.
Education and Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- Minimum years of experience: Experience requirements correlate with the internal job level.
- Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
Compensation
- Annual salary: $305,000–$385,000 USD.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
More jobs at Anthropic
Staff+ Software Engineer, Claude Science
Anthropic · San Francisco, United States
USD 405,000-485,000 per year
Head of Technical Training
Anthropic · San Francisco, United States
USD 0 per year
Applied AI, Research Engineer
Anthropic · New York City, United States, San Francisco, United States
USD 300,000-400,000 per year
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
TPM Manager, Infrastructure
Anthropic · New York City, United States, San Francisco, United States
USD 365,000-565,000 per year
Similar jobs
Lead Data Scientist, Platform Product
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 285,000-380,000 per year
Data Scientist, Product
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 285,000-380,000 per year
Member of Technical Staff (Software Engineer, Applied AI)
Perplexity AI · New York City, United States, Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year
AI Engineer, GTM Claudification
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Data Scientist, GTM
Anthropic · New York City, United States, San Francisco, United States
USD 285,000-380,000 per year
Full Stack Engineer, Education Labs
Anthropic · New York City, United States, San Francisco, United States
USD 300,000-405,000 per year
Senior Machine Learning Engineer, Trust
Airbnb · San Francisco, United States
USD 200,000-235,000 per year
Senior Machine Learning Engineer, Relevance and Personalization (Query Intelligence)
Airbnb · United States
USD 200,000-235,000 per year