Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 6
LLM
Machine Learning @ 3
Observability @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s Takeoff Intel team measures how much AI development is becoming AI-assisted. The team designs evaluations of AI research and development capabilities, builds internal telemetry, and develops quantitative methods to assess capability growth and support AI safety and situational awareness.
This generalist Research Engineer role focuses on evaluation infrastructure, large-scale data processing, and analysis tooling. The team is hiring at junior and senior levels.
Responsibilities
- Design, build, and run capability evaluations and measurement instruments at scale.
- Build data and analysis pipelines that transform large volumes of model outputs and telemetry into reliable metrics.
- Prototype new instruments quickly, validate them, and determine what to retain.
- Review and supervise AI-written code as part of the normal workflow.
- Collaborate with research scientists and partner teams to define what should be measured.
- Contribute to internal write-ups and public reporting.
Requirements
- Experience shipping an evaluation, data product, or research library end to end.
- Ability to prototype quickly and discard code when appropriate.
- Comfort handling messy, large-volume data without over-engineering.
- Experience running experiments on large language models, rather than only moving their outputs.
- Ability to work from vague research questions rather than detailed specifications.
- Strong communication skills and ability to collaborate closely with researchers.
Strong candidates may also have:
- Experience building evaluation harnesses or benchmark infrastructure for large language models.
- Experience with large-scale machine learning or data infrastructure, including self-driving, observability, or similar systems, alongside machine learning exposure.
- Experience building tools or libraries used by other researchers.
- A track record of identifying errors in AI-written code.
Examples of Work
- Anthropic ECI, an adaptation of the Epoch Capabilities Index used in recent system cards to measure capability acceleration.
- AI research and development capability assessments in Claude system cards.
- Data and research supporting When AI Builds Itself.
Education and Experience
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role, demonstrated through coursework, training, or professional experience.
- Required years of experience vary according to the internal job level.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.
More jobs at Anthropic
Applied AI Architect, Partnerships
Anthropic · London, United Kingdom
GBP 150,000-190,000 per year
Staff + Senior Software Engineer, Scaling
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Manager, Technical Deployment
Anthropic · London, United Kingdom
GBP 255,000-325,000 per year
Staff + Senior Software Engineer, Cloud Inference Launch Engineering
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Applied AI Engineer, DNB
Anthropic · London, United Kingdom
GBP 225,000-255,000 per year
Similar jobs
Staff Software Engineer, AI Reliability
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 325,000-485,000 per year
AI Engineer, GTM Claudification
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Staff Software Engineer, Machine Learning Platform
Stripe · South San Francisco, United States, Seattle, United States
USD 224,000-336,000 per year
Research Engineer, Model Evaluations
Anthropic · New York City, United States, San Francisco, United States
USD 500,000-850,000 per year
Distinguished Engineer, Core DevOps
GitLab · Canada, United Kingdom, United States
USD 250,000-349,000 per year
Senior Product Manager - AI/ML
ClickHouse · Argentina, Brazil, Canada, Chile, Colombia, Mexico, Peru, United States
USD 170,000-242,000 per year
Distinguished Engineer, Agentic SDLC & Non-Linear Productivity
GitLab · Canada, United Kingdom, Poland, United States
USD 250,000-349,000 per year
Senior Data Scientist, Ads Integrity
Reddit · United States
USD 190,800-267,100 per year