Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
Algorithms @ 4
CUDA
Debugging @ 6
Deep Learning @ 4
GPU @ 6
JAX @ 4
LLVM @ 4
Mentoring @ 4
Performance Analysis @ 6
PyTorch @ 4
Python @ 6
Robotics
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI to define the next era of computing, with GPUs powering computers, robots, and self-driving cars.
NVIDIA is hiring a software engineer for its Deep Learning and AI Compiler (DLC) team. The DLC powers deep learning inference across data centers, personal devices, automotive, and robotics. The compiler delivers inference performance, fast build times, reduced memory footprints, and ease of use through both Ahead-of-Time and Just-in-Time compilation.
Responsibilities
- Analyze deep learning networks and develop compiler optimization algorithms.
- Use CUDA to analyze and debug performance bottlenecks on GPUs.
- Define public APIs, develop performance optimizations and analysis, and implement compiler techniques for AI workloads and future NVIDIA GPUs.
Requirements
- Bachelor's, master's, or Ph.D. degree in Computer Science, Computer Engineering, a related field, or equivalent experience.
- 3+ years of relevant work or research experience in performance analysis and compiler optimizations.
- Experience with compiler technologies such as MLIR, LLVM, XLA, or Triton.
- Excellent C/C++ and Python programming and software design skills, including debugging, performance analysis, and test design.
- Ability to work independently, define project goals and scope, and lead development efforts.
- Strong interpersonal skills and the ability to work in a dynamic, product-oriented team.
Preferred Qualifications
- Proficiency in CPU and/or GPU architecture, especially modern NVIDIA GPUs such as Hopper and Blackwell.
- Understanding of deep learning models, algorithms, and frameworks such as PyTorch and JAX.
- Experience authoring GPU kernels and performing analysis using tools such as Nsight Compute.
- Experience mentoring early-career engineers and interns.
- Experience with new hardware bring-up.
Benefits
The position offers equity and benefits in addition to the base salary. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications for this job will be accepted at least until July 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.