Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 6
Debugging @ 4
Distributed Systems @ 4
JAX @ 4
LLM @ 7
Machine Learning @ 7
Networking
Observability @ 7
Performance Optimization
PyTorch @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic's ML Performance and Scaling team trains production pretrained models. This role works across the production training stack, including performance optimization, hardware debugging, experimental design, reliability, observability, and launch coordination. The role combines research and engineering, with an approximately 50/50 split between the two.
Responsibilities
- Own critical aspects of the production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
- Debug and resolve complex issues across the full stack, from hardware errors and networking to training dynamics and evaluation infrastructure.
- Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.
- Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams.
- Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.
- Add new capabilities to the training codebase, such as long-context support and novel architectures.
- Collaborate with teammates across San Francisco and London, as well as with the Tokens, Architectures, and Systems teams.
- Document systems, debugging approaches, and lessons learned to contribute to the team's institutional knowledge.
Requirements
- Hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
- Interest in both research and engineering work.
- Willingness to participate in on-call support for production systems, work extended hours during launches, and solve complex problems under pressure.
- Ability to debug complex, ambiguous problems across multiple layers of the stack.
- Clear communication and effective collaboration, including across time zones and during high-stress incidents.
- Passion for research engineering and continuous improvement of technical craft.
- Interest in the societal impacts of AI and responsible scaling.
- A bachelor's degree or equivalent combination of education, training, and experience.
- A field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
Strong candidates may also have experience training LLMs or working extensively with JAX, TPU, PyTorch, or other machine learning frameworks at scale; contributions to open-source LLM frameworks such as open_lm, llm-foundry, or mesh-transformer-jax; published research on model training, scaling laws, or ML systems; experience with production ML systems, observability tools, or evaluation infrastructure; or a background as a systems engineer, quant, or in another role requiring technical depth and operational excellence.
The role is highly operational and may involve responding to incidents on evenings and weekends during launches. The position requires working in the office five days per week in London.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Anthropic sponsors visas and states that it will make every reasonable effort to obtain a visa for candidates who receive an offer, with assistance from an immigration lawyer.