Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Deep Learning @ 7
GPU
HPC
JAX @ 4
LLM @ 4
Performance Analysis @ 4
Performance Optimization @ 4
PyTorch @ 4
Python @ 7
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's high-performance computing platforms power the AI revolution across many applications and industries. Within NVIDIA's software stack, CUTLASS is a popular open-source ecosystem dedicated to high-performance linear algebra and Tensor Core primitives. Since 2017, it has provided the community with C++ and Python abstractions for implementing custom matrix multiply (GEMM) and related mathematical and deep learning computations on NVIDIA GPUs.
The CUTLASS team is seeking an engineer enthusiastic about performance and eager to help bridge the gap between current performance and what is theoretically possible.
Responsibilities
- Benchmark the performance of state-of-the-art deep learning models' inference and training passes to identify key GPU kernel and fusion opportunities.
- Identify gaps between theoretical and realized performance, and suggest software improvements or model adjustments to resolve them.
- Develop tooling to automate the benchmarking, analysis, and performance optimization loop to push the limits of CUTLASS kernel performance within deep learning networks.
- Serve as the authoritative resource on kernel performance within the team.
- Engage with teams across NVIDIA, including GPU architecture, deep learning frameworks, and QA, as the performance representative for the CUTLASS team.
Requirements
- Master's or PhD degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 3+ years of relevant industry experience.
- Strong programming skills in Python and C++.
- Experience in software performance analysis and optimization.
- Deep understanding of computer architecture and familiarity with GPUs or similar parallel-processing architectures.
Preferred Qualifications
- Deep understanding of state-of-the-art deep learning model architectures.
- Hands-on experience benchmarking deep learning frameworks such as PyTorch, JAX, SGLang, vLLM, TRT-LLM, or others.
- Experience developing performance models and performance regression systems.
Compensation And Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Base salary is determined by location, experience, and the pay of employees in similar positions. Employees are also eligible for equity and benefits.
Applications will be accepted at least until June 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.