Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
Communication @ 6
Deep Learning @ 3
GPU @ 1
LLVM @ 4
Parallel Programming @ 4
Profiling
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an experienced Compiler Optimization Engineer to join its Compute Compiler Team. The team delivers features and improvements to CUDA and other compute compilers to help realize the potential of NVIDIA GPUs across deep learning, scientific computation, self-driving cars, and other workloads. The role focuses on a core compiler component for accelerating general-purpose computation on GPUs and involves working with compiler, hardware, and application teams.
Responsibilities
- Analyze the performance of application code running on NVIDIA GPUs using profiling tools.
- Identify opportunities for performance improvements in the MLIR/LLVM-based compiler middle-end optimizer.
- Identify and implement novel ideas in the compilation pipeline to achieve best-in-class performance for AI workloads.
- Design and develop compiler passes and optimizations that produce robust, maintainable, and supportable compiler tools.
- Interact with the open-source LLVM community to ensure tighter integration.
- Work with geographically distributed compiler, hardware, and application teams to oversee improvements and resolve problems.
- Contribute to deep-learning compiler technology spanning architecture design and support through higher-level languages.
Requirements
- Bachelor's, master's, or Ph.D. degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 8+ years of experience in compiler optimizations, such as loop, inter-procedural, and global optimizations.
- Excellent hands-on C++ programming skills.
- Understanding of a processor instruction set architecture; GPU ISA experience is a plus.
- Strong software engineering background, with a focus on developing robust and maintainable solutions to challenging problems.
- Good communication and documentation skills.
- Self-motivated approach to work.
Preferred Qualifications
- Master's or Ph.D. degree.
- Experience developing applications in CUDA or another parallel programming language.
- Deep understanding of parallel programming concepts.
- Experience with MLIR, LLVM, and/or Clang compiler development.
- Familiarity with deep learning frameworks and NVIDIA GPUs.
Benefits
The position offers equity and benefits, along with a comprehensive benefits package. NVIDIA states that the base salary is determined by location, experience, and the pay of employees in similar positions. Applications will be accepted at least until July 24, 2026. NVIDIA is an equal opportunity employer.