Senior Performance Compiler Engineer - Triton

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Algorithms @ 4 CUDA @ 6 Debugging @ 7 Deep Learning @ 4 GPU @ 4 HPC LLM Machine Learning @ 4 OpenCL @ 4 Parallel Programming @ 4 Performance Analysis @ 7 Python Software Development @ 7

Details

Role Overview

NVIDIA is looking for a Senior Performance Compiler Engineer to join its team and work on the open-source Triton compiler project. The role focuses on using compilers to improve AI performance on NVIDIA GPUs, enabling breakthroughs in large language models, agents, and other high-impact AI applications.

Responsibilities

  • Investigate the latest and future NVIDIA GPU hardware architecture and programming models.
  • Work on AI optimization opportunities by understanding advanced algorithms (like attention sinks and MoEs) and numerics (like block-scaled floating point).
  • Design and implement compiler technology using MLIR to optimize high-level kernel descriptions written in Triton’s Python DSL, with a focus on generating efficient, low-level GPU code.
  • When vital, use inline PTX to hand-tune critical code paths and extract peak performance from the hardware.
  • Perform an iterative optimization process—sometimes starting with the kernel, sometimes with the compiler—to find the most efficient path to peak performance.
  • Collaborate with teams across NVIDIA, including hardware architects and the CUDA compiler team, to influence future products and ensure maximum efficiency.

Requirements

  • Bachelor, Masters, or Ph.D. degree (or equivalent experience) in Computer Science, Computer Engineering, Applied Math, or a related field.
  • 8+ years of relevant industry experience in software development.
  • Strong C++ programming and software design skills, with an emphasis on performance analysis and debugging.
  • Experience in parallel programming, including CUDA/OpenCL GPU programming or other parallel models such as OpenMP.
  • Solid understanding of computer architecture and hands-on experience with assembly-level programming.

Ways to Stand Out

  • Experience tuning BLAS or deep learning library kernels.
  • Background in numerics and linear algebra.
  • Experience with machine learning compilers like TVM or MLIR.
  • Contributions to open-source projects, especially in AI/ML or compiler space.
  • Familiarity with the latest research in AI algorithms and numerics, and a strong track record of contributions to open-source projects, particularly in AI/ML, compiler, or high-performance computing domains.

Compensation and Benefits

  • Base salary range: 184,000 USD - 287,500 USD.
  • Eligible for equity and benefits (benefits link provided in the posting).

Application Notes

  • Applications accepted at least until May 12, 2026.
  • This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs