Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 6
LLM
Machine Learning @ 3
Observability @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s Takeoff Intel team measures how much AI development is becoming AI-assisted. The team designs evaluations of AI research and development capabilities, builds internal telemetry, and develops quantitative methods to assess capability growth and support AI safety and situational awareness.
This generalist Research Engineer role focuses on evaluation infrastructure, large-scale data processing, and analysis tooling. The team is hiring at junior and senior levels.
Responsibilities
- Design, build, and run capability evaluations and measurement instruments at scale.
- Build data and analysis pipelines that transform large volumes of model outputs and telemetry into reliable metrics.
- Prototype new instruments quickly, validate them, and determine what to retain.
- Review and supervise AI-written code as part of the normal workflow.
- Collaborate with research scientists and partner teams to define what should be measured.
- Contribute to internal write-ups and public reporting.
Requirements
- Experience shipping an evaluation, data product, or research library end to end.
- Ability to prototype quickly and discard code when appropriate.
- Comfort handling messy, large-volume data without over-engineering.
- Experience running experiments on large language models, rather than only moving their outputs.
- Ability to work from vague research questions rather than detailed specifications.
- Strong communication skills and ability to collaborate closely with researchers.
Strong candidates may also have:
- Experience building evaluation harnesses or benchmark infrastructure for large language models.
- Experience with large-scale machine learning or data infrastructure, including self-driving, observability, or similar systems, alongside machine learning exposure.
- Experience building tools or libraries used by other researchers.
- A track record of identifying errors in AI-written code.
Examples of Work
- Anthropic ECI, an adaptation of the Epoch Capabilities Index used in recent system cards to measure capability acceleration.
- AI research and development capability assessments in Claude system cards.
- Data and research supporting When AI Builds Itself.
Education and Experience
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role, demonstrated through coursework, training, or professional experience.
- Required years of experience vary according to the internal job level.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration.
More jobs at Anthropic
Salesforce Developer
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 270,000-345,000 per year
Business Systems Analyst, New Product Introduction
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 270,000-315,000 per year
Lead, Security Controls Assurance - SOX
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 410,000-510,000 per year
Lead Technical Instructor
Anthropic · New York City, United States
USD 270,000-310,000 per year
Product Designer, Evals & Prompts
Anthropic · San Francisco, United States
USD 305,000-385,000 per year
Similar jobs
Staff Software Engineer, AI Reliability
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 325,000-485,000 per year
AI Engineer, GTM Claudification
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Staff Software Engineer, Machine Learning Platform
Stripe · South San Francisco, United States, Seattle, United States
USD 224,000-336,000 per year
Research Engineer, Model Evaluations
Anthropic · New York City, United States, San Francisco, United States
USD 500,000-850,000 per year
Senior Data Scientist, Ads Integrity
Reddit · United States
USD 190,800-267,100 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Data Management Professional - Data Engineering - Corporate Actions
Bloomberg · Princeton, United States
USD 110,000-190,000 per year
Senior Manager, AI Software Engineering
SentinelOne · United States
USD 200,000-275,000 per year