Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Algorithms @ 3
CUDA @ 6
Debugging @ 7
Deep Learning @ 4
GPU @ 4
HPC
LLM
Machine Learning @ 4
Mathematics @ 4
OpenCL @ 4
Parallel Programming @ 4
Performance Analysis @ 7
Python
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Performance Compiler Engineer to join the open-source Triton compiler project. The role focuses on using compiler technology and new GPU technologies to improve AI performance on NVIDIA GPUs, enabling advances in large language models, agents, and other AI applications across training and inference.
Responsibilities
- Investigate current and future NVIDIA GPU hardware architectures and programming models.
- Analyze advanced AI algorithms, including attention sinks and mixture-of-experts (MoE), as well as numerical formats such as block-scaled floating point, to identify optimization opportunities.
- Design and implement compiler technology using MLIR to optimize high-level kernel descriptions written in Triton's Python DSL and generate efficient low-level GPU code.
- Use inline PTX when necessary to hand-tune critical code paths and maximize hardware performance.
- Iteratively optimize kernels and compiler technology to achieve peak performance.
- Collaborate with NVIDIA teams, including hardware architects and the CUDA compiler team, to influence future products and maximize efficiency.
Requirements
- Bachelor's, master's, or Ph.D. degree, or equivalent experience, in Computer Science, Computer Engineering, Applied Mathematics, or a related field.
- 8 or more years of relevant industry experience in software development.
- Strong C++ programming and software design skills, with an emphasis on performance analysis and debugging.
- Experience with parallel programming, including CUDA, OpenCL GPU programming, or other parallel models such as OpenMP.
- Solid understanding of computer architecture and hands-on experience with assembly-level programming.
Preferred Qualifications
- Experience tuning BLAS or deep learning library kernels.
- Background in numerics and linear algebra.
- Experience with machine learning compilers such as TVM or MLIR.
- Contributions to open-source projects, particularly in AI/ML or compiler-related areas.
- Familiarity with current research in AI algorithms and numerics.
- A strong track record of contributions to open-source projects in AI/ML, compiler, or high-performance computing domains.
Benefits
- Competitive salary with a base salary range of USD 184,000 to USD 287,500 per year.
- Eligibility for equity and benefits.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
Applications will be accepted at least until May 12, 2026. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Localization and Planning Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Customer Success Insights Engineer
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Relational Foundation Model Engineer, Modern Data Stack
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Insider Threat Detection Engineer
Nvidia · United States
USD 168,000-310,500 per year
Similar jobs
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Math Libraries Engineer - Sparsity in AI
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Deep Learning Algorithm Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year