Product Designer, Evals & Prompts

USD 305,000-385,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

A/B Testing @ 3 LLM @ 3 Python @ 3

Details

Anthropic is seeking a Product Designer, Evals & Prompts to join the Product Prompt and Eval Design team. The team designs prompts and evaluations for Claude's product surfaces, ensuring that the model and product remain aligned with user expectations, product strategy, and safety requirements across model launches.

This role focuses on building evaluations, the harness that runs them, and tools that enable designers to perform evaluation work independently. The role works with product surface owners, engineers, and prompt engineers during model releases.

Responsibilities

  • Write and revise prompts behind Claude's tools, features, and behaviors on product surfaces.
  • Test product surfaces, turn findings into prompt fixes, ship them, and confirm that users receive the intended prompts.
  • Build graders that validate prompt fixes and rerun evaluations on subsequent models.
  • Convert designers' manual rubrics into automated evaluations and review transcripts to identify gaps in evaluation coverage.
  • Build visual, low-code evaluation tools for non-engineering users, including tools to assemble comparison sets from real transcripts, convert plain-English rubrics into graders, compare prompt variants across models, and review results without a notebook.
  • Observe designers using evaluation tools and simplify the workflows.
  • Support model releases by testing each surface against new models, writing prompt fixes and migrations, and creating prompts for features launching with each model.
  • Establish and scale the evaluation harness, including a test environment for 50 to 100 tools with trustworthy settings.
  • Maintain evaluations across models and determine whether regressions originate in the harness or the model.
  • Package behaviors that prompting cannot fix for training, including graders, human-feedback questions, and good/bad preference pairs for qualities such as writing quality.

Requirements

  • Production-quality Python experience.
  • Experience building and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, regression suites, and the infrastructure to run evaluations across models.
  • Experience building internal tools with interfaces for people who do not write code.
  • Experience establishing test harnesses, sandboxing tool calls, and pinning settings to ensure comparable runs.
  • Experience shipping prompts or working closely with prompt developers, with an understanding of why a prompt that works on one model may fail on another.
  • Ability to read transcripts in addition to reviewing scores.

Preferred Qualifications

  • Experience working within a model-launch cycle.
  • A/B testing experience and the ability to connect offline evaluations to online outcomes.
  • Front-end or notebook-to-application experience, with an understanding of how to make evaluation results immediately legible.
  • Experience turning product rubrics into training signals, including graders, human-feedback questions, or preference pairs.
  • Care about how Claude behaves for users, not only whether a metric changes.

Education and Logistics

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and experience.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
  • Minimum years of experience: Experience requirements correlate with the internal job level.
  • Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.

Compensation

  • Annual salary: $305,000–$385,000 USD.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs