Performance Engineer, GPU

USD 280,000-850,000 per year
MIDDLE
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Algorithms CUDA @ 3 Communication @ 6 Distributed Systems @ 6 GPU @ 6 JAX @ 3 LLM Machine Learning @ 6 NCCL @ 3 NVLink @ 3 Profiling @ 3 PyTorch @ 3

Details

Anthropic is seeking a GPU Performance Engineer to architect and implement foundational systems that power Claude and advance the performance of large language models. The role focuses on maximizing GPU utilization and performance at scale, developing optimizations that enable new model capabilities and improve inference efficiency.

Working across hardware and software, the engineer will contribute to custom kernel development, distributed system architectures, low-level tensor core optimizations, and orchestration of large GPU clusters. Strong candidates will have a track record of delivering significant GPU performance improvements in production machine learning systems.

Responsibilities

  • Co-design attention mechanisms and algorithms for next-generation hardware architectures.
  • Develop custom kernels for emerging quantization formats and mixed-precision techniques.
  • Design distributed communication strategies for multi-node GPU clusters.
  • Optimize end-to-end training and inference pipelines for frontier language models.
  • Build performance modeling frameworks to predict and optimize GPU utilization.
  • Implement kernel fusion strategies to minimize memory bandwidth bottlenecks.
  • Create resilient systems for large-scale distributed training infrastructure.
  • Profile and eliminate performance bottlenecks in production serving infrastructure.
  • Partner with hardware vendors to influence future accelerator capabilities and software stacks.
  • Collaborate with researchers and engineers in an ambiguous, impact-driven environment.

Requirements

  • Deep experience with GPU programming and optimization at scale.
  • Ability to navigate complex systems from hardware interfaces to high-level machine learning frameworks.
  • Experience with one or more of the following areas:
    • GPU kernel development: CUDA, Triton, CUTLASS, FlashAttention, tensor core optimization.
    • ML compilers and frameworks: PyTorch or JAX internals, torch.compile, XLA, custom operators.
    • Performance engineering: kernel fusion, memory bandwidth optimization, and profiling with Nsight.
    • Distributed systems: NCCL, NVLink, collective communication, and model parallelism.
    • Low precision: INT8/FP8 quantization and mixed-precision techniques.
    • Production systems: large-scale training infrastructure, fault tolerance, and cluster orchestration.
  • Bachelor's degree or an equivalent combination of education, training, and experience.
  • Education or experience in a field relevant to the role.
  • Strong communication and collaborative problem-solving skills.

Work Policy

Staff are currently expected to work from one of Anthropic's offices at least 25% of the time. Some roles may require more time in the office.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration.

Additional Information

Applications are reviewed on a rolling basis, with no application deadline. Anthropic sponsors visas when possible and makes reasonable efforts to obtain visas for candidates who receive an offer.

More jobs at Anthropic

Similar jobs