Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 9
Agile @ 4
CUDA @ 4
Deep Learning @ 4
Engineering Management @ 7
GPU @ 4
GenAI
Generative AI
LLM @ 6
Leadership @ 7
NCCL @ 7
Performance Optimization @ 4
Profiling @ 6
PyTorch @ 6
Python @ 7
SGLang @ 6
Software Development @ 7
Technical Leadership @ 7
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an exceptional engineering manager to lead a world-class team advancing AI model deployment. The Deep Learning Inference team develops and optimizes open-source frameworks, including vLLM, SGLang, and FlashInfer, enabling scalable and efficient inference on NVIDIA GPUs for large language models, multimodal generative AI, and other applications across datacenter and edge environments.
Responsibilities
- Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
- Drive the strategy, roadmap, and execution of NVIDIA's inference framework engineering, with a focus on Client AI.
- Partner with compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
- Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
- Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications technologies including NIXL, NCCL, and NVSHMEM.
- Represent the team in roadmap and planning discussions and ensure alignment with NVIDIA's broader AI and software strategies.
- Foster a culture of technical excellence, open collaboration, and continuous innovation.
Requirements
- MS, PhD, or equivalent experience in Computer Science, Electrical or Computer Engineering, or a related field.
- At least 6 years of overall software development experience, including at least 3 years in technical leadership or engineering management.
- Strong background in C/C++ software design and development; Python proficiency is a plus.
- Hands-on experience with GPU programming using CUDA, Triton, and CUTLASS, as well as performance optimization.
- Proven experience deploying or optimizing deep learning models in production environments.
- Experience leading teams using Agile or collaborative software development practices.
Preferred Qualifications
- Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM.
- Deep understanding of multi-GPU communications using NIXL, NCCL, or NVSHMEM and distributed inference architectures.
- Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms.
- Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact.
- Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 2 and USD 224,000–356,500 for Level 3. Compensation will be determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 9, 2026. NVIDIA is an equal opportunity employer.