Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
CUDA @ 4
Debugging
Deep Learning @ 3
GPU @ 4
LLVM @ 6
Parallel Programming @ 4
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior AI Frameworks Engineer specializing in C++ and Python to help build the next frontier of the CUTLASS ecosystem: Pythonic CUTLASS (CUTLASS DSL). CUTLASS is an open-source ecosystem of high-performance math primitives that provides C++ template abstractions for implementing custom GEMM and related computations efficiently on NVIDIA GPUs. The initiative aims to bring high-performance execution and powerful abstractions into the Python environment, bridging low-level hardware primitives with high-level developer productivity.
Responsibilities
As a core contributor to the CUTLASS project, you will use systems programming and API design expertise to create a high-quality developer experience for GPU programming and kernel delivery.
- Design APIs that prioritize user productivity and provide a native feel for developers familiar with modern scientific computing and deep learning frameworks.
- Develop compilation infrastructure, including AST transformations and JIT-friendly execution, to lower Pythonic descriptions into high-performance GPU machine code.
- Create debugging tools, profiler integrations, and validation methodologies that simplify writing and using kernels.
- Build production-grade delivery infrastructure for the open-source community, including package distribution through wheels and conda, user-facing documentation, and testing.
Requirements
- MS or PhD degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 3 years of relevant experience.
- Strong proficiency in Python and C++, particularly in designing Python extensions and foreign function interfaces (FFI).
- Experience developing libraries or frameworks, with a focus on creating intuitive APIs for complex technical systems.
- Deep understanding of the Python ecosystem's delivery stack, including building, testing, and distributing high-performance compiled extensions.
Preferred Qualifications
- Active maintainer status or significant contributions to high-performance open-source libraries, AI frameworks, or compiler projects such as LLVM or MLIR.
- Understanding of compiler foundations, including intermediate representations (IR), lowering passes, or AST manipulation.
- Experience with GPU architecture and parallel programming models such as CUDA.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until June 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive work environment and equal opportunity employment.