Senior Deep Learning Software Engineer, Inference

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agile @ 1 Algorithms CUDA @ 4 Debugging @ 4 Deep Learning @ 1 GPU @ 4 GenAI Generative AI LLM NCCL @ 4 Profiling @ 4 PyTorch @ 6 Python @ 1 SGLang @ 6 Software Development @ 4 vLLM @ 6

Details

NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference. As a key contributor, you will help design, build, and optimize GPU-accelerated software that powers sophisticated AI applications. The team develops and maintains high-performance deep learning frameworks, including SGLang and vLLM, for efficient large-scale model serving and inference.

You will work closely with the deep learning community to implement the latest algorithms for public release in SGLang, vLLM, and other deep learning frameworks. Your work will focus on identifying and driving performance improvements for state-of-the-art large language models and generative AI models across NVIDIA accelerators, from datacenter GPUs to edge systems. You will use open-source tools and plugins, including CUTLASS, OpenAI Triton, NCCL, and CUDA kernels, to implement and optimize model-serving pipelines.

Responsibilities

  • Optimize, analyze, and tune deep learning models across domains including large language models, multimodal AI, and generative AI.
  • Scale deep learning model performance across different architectures and types of NVIDIA accelerators.
  • Contribute features and code to NVIDIA inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.
  • Collaborate with teams working on frameworks, NVIDIA libraries, and innovative inference-optimization solutions.

Requirements

  • Master's degree, PhD, or equivalent experience in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
  • Six or more years of relevant software development experience.
  • Excellent C/C++ programming and software design skills.
  • Software Agile experience is helpful; Python experience is a plus.
  • Experience training, deploying, or optimizing the inference of deep learning models in production is a plus.
  • Background in performance modeling, profiling, debugging, and code optimization, or architectural knowledge of CPUs and GPUs, is a plus.

Preferred Qualifications

  • Contributions to deep learning software projects such as PyTorch, vLLM, or SGLang.
  • Experience with multi-GPU communications, including NCCL or NVSHMEM.
  • Experience building and shipping products to enterprise customers.
  • GPU programming experience with CUDA, OpenAI Triton, or CUTLASS.

Compensation and Benefits

The base salary is determined by location, experience, and compensation for similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until September 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs