Research Engineer, Model Evaluations

USD 500,000-850,000 per year
MIDDLE
✅ Remote ✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 3 Communication @ 3 Data Pipelines @ 3 Data Visualization @ 3 Distributed Systems @ 3 LLM Machine Learning @ 3 Observability @ 3 Python @ 6 Reinforcement Learning Statistics @ 3

Details

Anthropic is looking for Research Engineers to build evaluations that characterize Claude's capabilities and personality. The role involves turning ambiguous notions of intelligence into clear, defensible metrics and building reliable infrastructure to run evaluations at scale. You will partner with research teams throughout the lifecycle of new capabilities, from defining what to measure through interpreting results during training.

Responsibilities

  • Design and run evaluations of Claude's reasoning, agentic behavior, knowledge, and safety properties, and produce visualizations for researchers and decision-makers.
  • Build and harden the distributed evaluation execution platform so that hundreds of evaluations run reliably against checkpoints during production reinforcement learning training runs.
  • Own dashboards used by researchers and leadership to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions difficult to miss.
  • Debug anomalous evaluation results during training runs and determine whether the cause is a model change or an infrastructure issue.
  • Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations.
  • Partner with research teams to define measurements and interpret results as training progresses.
  • Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks.
  • Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences.

Requirements

  • Strong Python programming skills, including production or research infrastructure.
  • Experience building or operating distributed systems, data pipelines, or other infrastructure that must be reliable at scale.
  • Clear written and verbal communication, especially when explaining technical results to non-specialists.
  • Comfort operating in an on-call or production-support capacity when training runs are live.
  • Care about the societal impacts of the work and an interest in steering powerful AI to be safe and beneficial.

Preferred Qualifications

  • Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding.
  • Background in data visualization and experience building trusted, usable dashboards.
  • Experience developing robust evaluation metrics for language models.
  • Experience with observability, monitoring, or experiment-tracking systems.
  • Background in statistics and experimental design.
  • Experience with large-scale dataset sourcing, curation, and processing.
  • Experience running or supporting machine learning training infrastructure.
  • Flexibility to operate across team boundaries and a willingness to pick up additional work.
  • Enjoyment of pair programming.

Representative Projects

  • Stand up a new evaluation for a specific reasoning capability by defining the task, building the dataset, implementing scoring, validating against known signals, and shipping a dashboard.
  • Diagnose a mid-training regression and determine whether the cause is the model, evaluation harness, data, or infrastructure.
  • Improve a flaky distributed evaluation pipeline through better retries, observability, and faster feedback.
  • Partner with a research team to define what good performance means in a new capability area and translate it into measurable artifacts.

Education and Logistics

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
  • Minimum years of experience: Experience requirements correlate with the internal job level.
  • Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
  • Anthropic sponsors visas and makes reasonable efforts to support visa applications, with assistance from an immigration lawyer.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

More jobs at Anthropic

Similar jobs