Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 4
Deep Learning @ 7
GPU @ 7
HPC
JAX
LLM @ 3
Machine Learning @ 6
Performance Analysis @ 4
Profiling @ 4
PyTorch
Python @ 7
TensorRT
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are now seeking a Senior Deep Learning Performance Architect. NVIDIA is looking for outstanding Performance Architects with a background in performance analysis, performance modeling, and AI/deep learning to help analyze and develop the next generation of architectures that accelerate AI and high-performance computing applications.
Responsibilities
- Develop innovative architectures to extend the state of the art in deep learning performance and efficiency
- Analyze performance, cost and power trade-offs by developing analytical models, simulators and test suites
- Understand and analyze the interplay of hardware and software architectures on future algorithms, programming models and applications
- Evaluate PPA (performance, power, area) for hardware features and system level architectural trade-offs. Develop high level simulators in C++/Python
- Actively collaborate with software, product and research teams to guide the direction of deep learning HW and SW
Requirements
- MS or PhD in Computer Science, Computer Engineering, Electrical Engineering or equivalent experience
- 6+ years of relevant meaningful work experience
- Strong background in GPU or Deep Learning ASIC architecture for distributed training and/or inference spanning multi-chip/multi-node
- Experience with performance modeling, architecture simulation, profiling, and analysis
- Solid foundation in machine learning and deep learning. Understanding of modern transformer-based architectures and their performance at scale.
- Strong programming skills in Python, C, C++
Ways to stand out from the crowd
- Background with deep neural network training, inference and optimization in leading frameworks (e.g. Pytorch, JAX, TensorRT)
- Familiarity with advanced optimizations and SW/HW co-design in LLM training and inference
- Exposure to using AI to accelerate SW engineering
- Demonstration of self-motivation and creative / critical thinking
NVIDIA describes this as part of its Deep Learning Architecture team, helping build real-time, efficient computing platforms for its success in this rapidly growing field.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
More jobs at Nvidia
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Director, Autonomous Vehicles Platform
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Developer Technology Engineer - Edge Agentic Ai
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - Halos
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Drive Os Communication Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Deep Learning Performance Architect
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, AI Stack And Performance - DGX Station
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Ai Networking
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Tools Engineer – CUDA Tile
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year