Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
Debugging @ 7
Deep Learning
GPU
HPC
OpenCL @ 4
Python
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's high-performance computing platforms are powering the AI revolution across many applications and industries. Within NVIDIA's software stack, CUTLASS is an open-source ecosystem dedicated to high-performance linear algebra and Tensor Core primitives. Since 2017, it has provided C++ and Python abstractions for implementing custom matrix multiply (GEMM) and related mathematical and deep learning computations on NVIDIA GPUs.
The CUTLASS team develops and optimizes math kernels to extract the highest performance from NVIDIA hardware architectures.
Responsibilities
- Write Tensor Core-based deep learning kernels, including grouped-GEMM, attention, and convolution, using CUTLASS CUDA C++ and Python DSL for Blackwell, Rubin, and future architectures.
- Optimize kernels for peak throughput on both silicon and software performance simulators.
- Collaborate with NVIDIA teams across GPU architecture, NVVM/PTX compiler, CUDA libraries, and deep learning frameworks to ensure fast, functional, and timely kernel delivery to customers.
Requirements
- Master's or PhD degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 3+ years of relevant industry experience.
- Strong proficiency in C++ programming and software design, including debugging, performance evaluation, and testing.
- Experience with CUDA, OpenCL, HIP, SYCL, Mojo, Pallas, Triton, Mosaic, Halide, or another general-purpose or domain-specific programming language targeting highly parallel accelerators.
- Deep understanding of computer architecture and some experience working at the assembly level.
Preferred Qualifications
- Experience writing code specifically targeting NVIDIA Tensor Cores, particularly through PTX or CUDA/cuTile.
- Open-source contributions to math kernel libraries or frameworks.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications for this job will be accepted at least until June 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior DevTech Compute Engineer, Compression and Data Processing
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Performance Compiler Engineer - Triton
Nvidia · Redmond, United States
USD 184,000-287,500 per year
Senior AI Compiler Engineer, MLIR
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Deep Learning Compiler Engineer - XLA
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUTLASS Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Simulation Architect
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year