Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 7
CUDA @ 4
Deep Learning @ 6
GPU @ 6
LLM
Microservices
OpenCL @ 4
Performance Analysis @ 4
Performance Optimization @ 4
Profiling @ 4
PyTorch @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a senior engineer focused on performance analysis and optimization of deep learning workloads. The role works across the hardware and software stack, from GPU architecture to deep learning frameworks, to improve inference performance and influence future hardware and software roadmaps.
Responsibilities
- Implement language and multimodal model inference as part of NVIDIA Inference Microservices (NIMs).
- Contribute new features, fix bugs, and deliver production code to TRT-LLM, NVIDIA's open-source inference serving library.
- Profile and analyze bottlenecks across the full inference stack to improve inference performance.
- Benchmark state-of-the-art deep learning model inference offerings and perform competitive analysis of NVIDIA software and hardware stacks.
- Collaborate with software and hardware co-design teams to enable the creation of next-generation AI-powered services.
Requirements
- PhD in Computer Science, Electrical Engineering, Computer Science and Engineering, or equivalent experience.
- 5+ years of experience.
- Strong background in deep learning and neural networks, particularly inference.
- Experience with performance profiling, analysis, and optimization, especially for GPU-based applications.
- Proficiency in C++ and PyTorch or equivalent frameworks.
- Deep understanding of computer architecture and familiarity with GPU architecture fundamentals.
Preferred Qualifications
- Proven experience with processor- and system-level performance optimization.
- Deep understanding of modern large language model architectures.
- Strong fundamentals in algorithms.
- GPU programming experience with CUDA or OpenCL.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior DL Algorithms Engineer - Inference Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Deep Learning Compiler Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Systems Software Engineer, AI Stack and Performance - DGX Station
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal High-Performance LLM Training Engineer
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer - Python Numerical Computing Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Deep Learning Inference – TensorRT
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year