Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms @ 4
CUDA @ 6
Debugging @ 7
Deep Learning @ 4
GPU @ 4
HPC
LLM
Machine Learning @ 4
OpenCL @ 4
Parallel Programming @ 4
Performance Analysis @ 7
Python
Software Development @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Role Overview
NVIDIA is looking for a Senior Performance Compiler Engineer to join its team and work on the open-source Triton compiler project. The role focuses on using compilers to improve AI performance on NVIDIA GPUs, enabling breakthroughs in large language models, agents, and other high-impact AI applications.
Responsibilities
- Investigate the latest and future NVIDIA GPU hardware architecture and programming models.
- Work on AI optimization opportunities by understanding advanced algorithms (like attention sinks and MoEs) and numerics (like block-scaled floating point).
- Design and implement compiler technology using MLIR to optimize high-level kernel descriptions written in Triton’s Python DSL, with a focus on generating efficient, low-level GPU code.
- When vital, use inline PTX to hand-tune critical code paths and extract peak performance from the hardware.
- Perform an iterative optimization process—sometimes starting with the kernel, sometimes with the compiler—to find the most efficient path to peak performance.
- Collaborate with teams across NVIDIA, including hardware architects and the CUDA compiler team, to influence future products and ensure maximum efficiency.
Requirements
- Bachelor, Masters, or Ph.D. degree (or equivalent experience) in Computer Science, Computer Engineering, Applied Math, or a related field.
- 8+ years of relevant industry experience in software development.
- Strong C++ programming and software design skills, with an emphasis on performance analysis and debugging.
- Experience in parallel programming, including CUDA/OpenCL GPU programming or other parallel models such as OpenMP.
- Solid understanding of computer architecture and hands-on experience with assembly-level programming.
Ways to Stand Out
- Experience tuning BLAS or deep learning library kernels.
- Background in numerics and linear algebra.
- Experience with machine learning compilers like TVM or MLIR.
- Contributions to open-source projects, especially in AI/ML or compiler space.
- Familiarity with the latest research in AI algorithms and numerics, and a strong track record of contributions to open-source projects, particularly in AI/ML, compiler, or high-performance computing domains.
Compensation and Benefits
- Base salary range: 184,000 USD - 287,500 USD.
- Eligible for equity and benefits (benefits link provided in the posting).
Application Notes
- Applications accepted at least until May 12, 2026.
- This posting is for an existing vacancy.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Ai Networking
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Compiler Engineer, AI Inference Platforms
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Software Engineer, Deep Learning Inference - TensorRT
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Performance And Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Deep Learning Performance Architect
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year