Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 3
Data Pipelines @ 3
Data Visualization @ 3
Distributed Systems @ 3
LLM
Machine Learning @ 3
Observability @ 3
Python @ 6
Reinforcement Learning
Statistics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is looking for Research Engineers to build evaluations that characterize Claude's capabilities and personality. The role involves turning ambiguous notions of intelligence into clear, defensible metrics and building reliable infrastructure to run evaluations at scale. You will partner with research teams throughout the lifecycle of new capabilities, from defining what to measure through interpreting results during training.
Responsibilities
- Design and run evaluations of Claude's reasoning, agentic behavior, knowledge, and safety properties, and produce visualizations for researchers and decision-makers.
- Build and harden the distributed evaluation execution platform so that hundreds of evaluations run reliably against checkpoints during production reinforcement learning training runs.
- Own dashboards used by researchers and leadership to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions difficult to miss.
- Debug anomalous evaluation results during training runs and determine whether the cause is a model change or an infrastructure issue.
- Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations.
- Partner with research teams to define measurements and interpret results as training progresses.
- Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks.
- Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences.
Requirements
- Strong Python programming skills, including production or research infrastructure.
- Experience building or operating distributed systems, data pipelines, or other infrastructure that must be reliable at scale.
- Clear written and verbal communication, especially when explaining technical results to non-specialists.
- Comfort operating in an on-call or production-support capacity when training runs are live.
- Care about the societal impacts of the work and an interest in steering powerful AI to be safe and beneficial.
Preferred Qualifications
- Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding.
- Background in data visualization and experience building trusted, usable dashboards.
- Experience developing robust evaluation metrics for language models.
- Experience with observability, monitoring, or experiment-tracking systems.
- Background in statistics and experimental design.
- Experience with large-scale dataset sourcing, curation, and processing.
- Experience running or supporting machine learning training infrastructure.
- Flexibility to operate across team boundaries and a willingness to pick up additional work.
- Enjoyment of pair programming.
Representative Projects
- Stand up a new evaluation for a specific reasoning capability by defining the task, building the dataset, implementing scoring, validating against known signals, and shipping a dashboard.
- Diagnose a mid-training regression and determine whether the cause is the model, evaluation harness, data, or infrastructure.
- Improve a flaky distributed evaluation pipeline through better retries, observability, and faster feedback.
- Partner with a research team to define what good performance means in a new capability area and translate it into measurable artifacts.
Education and Logistics
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- Minimum years of experience: Experience requirements correlate with the internal job level.
- Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and makes reasonable efforts to support visa applications, with assistance from an immigration lawyer.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.
More jobs at Anthropic
Technical Program Manager, Silicon
Anthropic · San Francisco, United States, New York City, United States
USD 365,000-435,000 per year
Researcher, Cybersecurity Products
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Technical Program Manager, RL Research
Anthropic · San Francisco, United States, New York City, United States
USD 365,000-435,000 per year
Product Design Manager
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 385,000-460,000 per year
Insider Risk Investigator
Anthropic · Washington, United States, Boston, United States, New York City, United States
USD 245,000-305,000 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Research Engineer, Economic Research Data Platform
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Staff Advanced Analytics, Digital & AI Products
Airbnb · Bengaluru, India
INR 4,410,000-6,300,000 per year
Staff Software Engineer, Machine Learning Platform
Stripe · South San Francisco, United States, Seattle, United States
USD 224,000-336,000 per year
Staff Software Engineer, Environments Infrastructure
Anthropic · San Francisco, United States, New York City, United States
USD 405,000-605,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year