Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 4
CUDA @ 4
Deep Learning @ 7
GPU @ 7
HPC
JAX @ 6
Machine Learning @ 6
OpenCL @ 4
Profiling @ 4
PyTorch @ 6
Python @ 7
Robotics
TensorRT @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Deep Learning Performance Architect to help analyze and develop next-generation architectures that accelerate artificial intelligence and high-performance computing applications.
Responsibilities
- Develop innovative architectures to extend the state of the art in deep learning performance and efficiency.
- Analyze performance, cost, and power trade-offs by developing analytical models, simulators, and test suites.
- Understand and analyze the interplay of hardware and software architectures on future algorithms, programming models, and applications.
- Develop, analyze, and harness deep learning frameworks, libraries, and compilers.
- Collaborate with software, product, and research teams to guide the direction of deep learning hardware and software.
Requirements
- Master’s or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- 6+ years of meaningful work experience.
- Strong background in GPU or deep learning ASIC architecture for training and/or inference.
- Experience with performance modeling, architecture simulation, profiling, and analysis.
- Solid foundation in machine learning and deep learning.
- Strong programming skills in Python, C, and C++.
Preferred Qualifications
- Background with deep neural network training, inference, and optimization in leading frameworks such as PyTorch, JAX, and TensorRT.
- Experience with relevant libraries, compilers, and languages, including cuDNN, cuBLAS, CUTLASS, MLIR, Triton, CUDA, and OpenCL.
- Experience with the architecture of or workload analysis on other deep learning accelerators.
- Self-motivation, critical thinking, and the ability to think outside the box.
Additional Information
NVIDIA’s GPUs run artificial intelligence algorithms and support applications including robotics and self-driving cars. The role is part of NVIDIA’s Deep Learning Architecture team, focused on building real-time, efficient computing platforms.
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The employee will also be eligible for equity and benefits. Applications will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.