Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile @ 1
Algorithms
CUDA @ 4
Debugging @ 4
Deep Learning @ 1
GPU @ 4
GenAI
Generative AI
LLM
NCCL @ 4
Profiling @ 4
PyTorch @ 6
Python @ 1
SGLang @ 6
Software Development @ 4
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference. As a key contributor, you will help design, build, and optimize GPU-accelerated software that powers sophisticated AI applications. The team develops and maintains high-performance deep learning frameworks, including SGLang and vLLM, for efficient large-scale model serving and inference.
You will work closely with the deep learning community to implement the latest algorithms for public release in SGLang, vLLM, and other deep learning frameworks. Your work will focus on identifying and driving performance improvements for state-of-the-art large language models and generative AI models across NVIDIA accelerators, from datacenter GPUs to edge systems. You will use open-source tools and plugins, including CUTLASS, OpenAI Triton, NCCL, and CUDA kernels, to implement and optimize model-serving pipelines.
Responsibilities
- Optimize, analyze, and tune deep learning models across domains including large language models, multimodal AI, and generative AI.
- Scale deep learning model performance across different architectures and types of NVIDIA accelerators.
- Contribute features and code to NVIDIA inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.
- Collaborate with teams working on frameworks, NVIDIA libraries, and innovative inference-optimization solutions.
Requirements
- Master's degree, PhD, or equivalent experience in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
- Six or more years of relevant software development experience.
- Excellent C/C++ programming and software design skills.
- Software Agile experience is helpful; Python experience is a plus.
- Experience training, deploying, or optimizing the inference of deep learning models in production is a plus.
- Background in performance modeling, profiling, debugging, and code optimization, or architectural knowledge of CPUs and GPUs, is a plus.
Preferred Qualifications
- Contributions to deep learning software projects such as PyTorch, vLLM, or SGLang.
- Experience with multi-GPU communications, including NCCL or NVSHMEM.
- Experience building and shipping products to enterprise customers.
- GPU programming experience with CUDA, OpenAI Triton, or CUTLASS.
Compensation and Benefits
The base salary is determined by location, experience, and compensation for similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until September 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.