Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 9
Agile @ 4
CUDA @ 4
Deep Learning @ 9
Engineering Management @ 6
GPU @ 4
GenAI
Generative AI
LLM
Leadership @ 6
NCCL @ 7
OSS
Performance Optimization @ 4
Profiling
PyTorch
Python @ 7
SGLang
Software Development @ 6
Technical Leadership @ 6
TensorRT
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the software powering today’s most sophisticated AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs. The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible — including SGLang, vLLM, and FlashInfer. Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices.
What you’ll be doing
- Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
- Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.
- Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
- Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
- Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
- Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies.
- Foster a culture of technical excellence, open collaboration, and continuous innovation.
What we need to see
- MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
- 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.
- Strong background in C/C++ software design and development; proficiency in Python is a plus.
- Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
- Proven record of deploying or optimizing deep learning models in production environments.
- Experience leading teams using Agile or collaborative software development practices.
Ways to Stand out from The Crowd
- Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM.
- Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures.
- Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms.
- Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact.
- Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering.
Benefits
With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. You will also be eligible for equity and benefits.