Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA
Distributed Systems
GPU
Go @ 6
JAX
Java @ 6
LLM
Machine Learning
Performance Optimization
Prioritization @ 7
Profiling
PyTorch
Python @ 6
Rust @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's Interpretability team reverse-engineers trained models to develop a mechanistic understanding of how they work and improve the safety of advanced AI systems. This role focuses on building the engineering and infrastructure required to scale interpretability research and apply it to production safety audits.
Responsibilities
- Build and maintain specialized inference and training infrastructure for interpretability research, including instrumented forward and backward passes, activation extraction, and steering vector application.
- Resolve scaling and efficiency bottlenecks through profiling, optimization, and collaboration with infrastructure teams.
- Design tools, abstractions, and platforms that enable researchers to experiment rapidly without engineering barriers.
- Help bring interpretability research into production safety audits with real deadlines and high reliability expectations.
- Work across the stack, from model internals and accelerator-level optimization to user-facing research tooling.
Requirements
- 5–10+ years of experience building software.
- High proficiency in at least one programming language, such as Python, Rust, Go, or Java, and productive use of Python.
- Ability to learn unfamiliar technical domains quickly and investigate bottlenecks across different layers of the stack.
- Strong prioritization skills and comfort operating with ambiguity and questioning assumptions.
- Preference for fast-moving, collaborative projects.
- Interest in interpretability research and its role in AI safety; prior research experience is not required.
- Care for the societal impacts and ethics of technical work.
- Ability to work closely with researchers and translate research needs into engineering solutions.
- Bachelor's degree or an equivalent combination of education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Additional Experience
Strong candidates may also have experience with:
- Performance optimization of large-scale distributed systems.
- Language modeling fundamentals and transformers.
- High-performance LLM optimization, including memory management, compute efficiency, parallelism strategies, and inference throughput optimization.
- Mainstream machine learning stacks, such as PyTorch and CUDA on GPUs or JAX and XLA on TPUs.
- Research collaboration and research tooling, or direct research involving complex engineering challenges.
Representative Projects
- Building Garcon, a tool for instrumenting LLMs to extract internal activations.
- Designing and optimizing pipelines to collect and shuffle petabytes of transformer activations.
- Profiling and optimizing machine learning training jobs, including multi-GPU parallelism and memory optimization.
- Building steered inference systems that apply targeted interventions to model internals at scale.
Location and Work Policy
The role is based in Anthropic's San Francisco office, with exceptional candidates considered for remote work on a case-by-case basis. Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
Compensation
The annual salary range is $315,000–$560,000 USD.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office designed for collaboration. Anthropic sponsors visas and makes reasonable efforts to support visa applications, with assistance from an immigration lawyer.