Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile @ 1
Algorithms
CUDA @ 1
Debugging @ 4
Deep Learning @ 1
GPU @ 1
GenAI
Generative AI
LLM
NCCL @ 4
Profiling @ 4
PyTorch @ 6
Python @ 1
SGLang @ 6
Software Development @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA seeks a Senior Software Engineer specializing in deep learning inference. As a key contributor, you will help design, build, and optimize GPU-accelerated software powering sophisticated AI applications. The team develops and maintains high-performance open-source frameworks for efficient large-scale model serving and inference, facilitating the deployment and serving of language models.
You will work closely with the deep learning community to implement the latest algorithms for public release in inference frameworks. Your work will focus on identifying and driving performance improvements for state-of-the-art LLM and generative AI models across NVIDIA accelerators, from datacenter GPUs to edge SoCs. You will use open-source tools and plugins—including CUTLASS, OAI Triton, NCCL, and CUDA kernels—to implement and optimize model-serving pipelines.
Responsibilities
- Optimize, analyze, and tune deep learning models in domains including LLMs, multimodal AI, and generative AI.
- Scale deep learning model performance across different architectures and types of NVIDIA accelerators.
- Contribute features and code to NVIDIA inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.
- Collaborate with teams across frameworks, NVIDIA libraries, and inference-optimization initiatives.
Requirements
- Master's degree, PhD, or equivalent experience in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
- 5+ years of relevant software development experience.
- Excellent C/C++ programming and software design skills.
- Software Agile skills are helpful; Python experience is a plus.
- Experience training, deploying, or optimizing the inference of deep learning models in production is a plus.
- Background in performance modeling, profiling, debugging, and code optimization, or architectural knowledge of CPUs and GPUs, is a plus.
- GPU programming experience with CUDA, OAI Triton, or CUTLASS is a plus.
Preferred Qualifications
- Contributions to deep learning software projects such as PyTorch, vLLM, or SGLang.
- Experience with multi-GPU communications, including NCCL or NVSHMEM.
Compensation And Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Salary is determined based on location, experience, and compensation for employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.