Senior Deep Learning Software Engineer, Inference

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Agile @ 1 Algorithms CUDA @ 1 Debugging @ 4 Deep Learning @ 1 GPU @ 1 GenAI Generative AI LLM NCCL @ 4 Profiling @ 4 PyTorch @ 6 Python @ 1 SGLang @ 6 Software Development @ 6 vLLM @ 6

Details

NVIDIA seeks a Senior Software Engineer specializing in deep learning inference. As a key contributor, you will help design, build, and optimize GPU-accelerated software powering sophisticated AI applications. The team develops and maintains high-performance open-source frameworks for efficient large-scale model serving and inference, facilitating the deployment and serving of language models.

You will work closely with the deep learning community to implement the latest algorithms for public release in inference frameworks. Your work will focus on identifying and driving performance improvements for state-of-the-art LLM and generative AI models across NVIDIA accelerators, from datacenter GPUs to edge SoCs. You will use open-source tools and plugins—including CUTLASS, OAI Triton, NCCL, and CUDA kernels—to implement and optimize model-serving pipelines.

Responsibilities

  • Optimize, analyze, and tune deep learning models in domains including LLMs, multimodal AI, and generative AI.
  • Scale deep learning model performance across different architectures and types of NVIDIA accelerators.
  • Contribute features and code to NVIDIA inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.
  • Collaborate with teams across frameworks, NVIDIA libraries, and inference-optimization initiatives.

Requirements

  • Master's degree, PhD, or equivalent experience in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
  • 5+ years of relevant software development experience.
  • Excellent C/C++ programming and software design skills.
  • Software Agile skills are helpful; Python experience is a plus.
  • Experience training, deploying, or optimizing the inference of deep learning models in production is a plus.
  • Background in performance modeling, profiling, debugging, and code optimization, or architectural knowledge of CPUs and GPUs, is a plus.
  • GPU programming experience with CUDA, OAI Triton, or CUTLASS is a plus.

Preferred Qualifications

  • Contributions to deep learning software projects such as PyTorch, vLLM, or SGLang.
  • Experience with multi-GPU communications, including NCCL or NVSHMEM.

Compensation And Benefits

The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Salary is determined based on location, experience, and compensation for employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until July 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs