Research Engineer, Visual Knowledge Work

USD 350,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Agentic Systems Computer Vision @ 6 Deep Learning ETL LLM Machine Learning @ 6 Reinforcement Learning @ 3

Details

Anthropic is looking for a research engineer to help develop visual and spatial reasoning capabilities for large language models. The role focuses on creating training data and reinforcement learning environments for visual knowledge work, including identifying long-horizon and vision-heavy tasks, building evaluations, designing rewards, and scaling data. You will collaborate with external vendors, pretraining, reinforcement learning, and product teams to translate these environments into real-world knowledge work capabilities.

Responsibilities

  • Own the end-to-end data strategy for vision capabilities, including building evaluations and scaling reinforcement learning environments.
  • Manage technical relationships with external data vendors, including writing task specifications, evaluating visual data and annotation quality, and iterating on reward design.
  • Develop and improve quality assurance frameworks to detect reward hacking and ensure environment quality at scale.
  • Run generalization experiments to measure how data strategy changes affect multimodal capabilities on held-out evaluations.
  • Partner with pretraining, reinforcement learning, and product teams.
  • Conduct research to ensure teams are aligned on capability development.

Requirements

  • 7+ years of experience in machine learning, computer vision, and software engineering through industry, academia, or other projects.
  • Experience with reinforcement learning, reward design, or training data curation for large language or vision-language models.
  • Familiarity with the architecture, training, and operation of large vision-language models.
  • Comfortable managing technical vendor relationships and iterating quickly based on feedback.
  • Results-oriented, flexible, and focused on impact.
  • Awareness of the societal impacts of this work.
  • A bachelor's degree or equivalent combination of education, training, and experience in a relevant field.

Preferred Experience

  • Designing evaluations or benchmarks for large language models or vision-language models.
  • Large-scale pretraining, supervised learning, and reinforcement learning on language models.
  • Deep learning research involving images, video, or other modalities.
  • Developing complex agentic systems using large language models.
  • Large-scale ETL and data pipeline development.

Representative Projects

  • Writing vendor-facing specifications for visual reinforcement learning training tasks and iterating on coverage, quality, and reward design.
  • Running experiments to determine optimal training data mixes and parameters for a synthetically generated vision dataset.
  • Fine-tuning Claude to maximize performance using a particular set of agent tools or skills.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office environment for collaboration.

More jobs at Anthropic

Similar jobs