Senior Software Engineer - AI Inference

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 7 Communication @ 7 Distributed Systems @ 3 GPU @ 4 InfiniBand @ 6 LLM @ 4 NCCL @ 6 NVLink @ 6 Profiling @ 4 PyTorch @ 6 Python @ 7 SGLang @ 4 SRE @ 7 vLLM @ 4

Details

NVIDIA is seeking a Senior Software Engineer - AI Inference to advance open-source LLM serving by contributing directly to upstream inference engines such as vLLM and SGLang. The role focuses on ensuring these engines run effectively on NVIDIA GPUs and systems while improving the underlying stack for high-throughput, low-latency inference at scale.

This is a hands-on position for an engineer who enjoys investigating performance bottlenecks, designing runtime improvements, and delivering high-quality changes for open-source communities and production deployments.

Responsibilities

  • Contribute features, fixes, and optimizations upstream to vLLM and SGLang by authoring pull requests, participating in reviews, writing benchmarks and tests, and helping drive designs to completion.
  • Implement and optimize inference-runtime capabilities, including batching and scheduling policies, streaming, request lifecycle management, and KV-cache efficiency through paging and sharding.
  • Profile and improve hot paths across layers, from Python orchestration to C++ and CUDA kernels, using data to guide optimization work.
  • Improve multi-GPU inference performance and reliability through parallelism strategies, communication patterns, and resource utilization across NVIDIA platforms.
  • Build and maintain performance and correctness regression tests to prevent slowdowns and ensure stable behavior across model and hardware configurations.
  • Collaborate with model, platform, and SRE teams to translate production requirements into upstreamable solutions with strong operability and maintainability.

Requirements

  • 5+ years of experience building production software, with solid systems engineering fundamentals and a track record of delivering performance or reliability improvements.
  • Experience with LLM inference and serving stacks such as vLLM or SGLang, including an understanding of the tradeoffs that drive production performance.
  • Strong programming skills in Python plus C++ and/or CUDA, with the ability to debug and optimize performance-critical code.
  • Experience with profiling and performance investigation, including microbenchmarks, flame graphs, and GPU profiling, along with a measurement-driven mindset.
  • Familiarity with distributed systems concepts and concurrency, including queues and schedulers, multi-process and multi-threaded systems, and scaling across GPUs and nodes.
  • Strong communication skills and comfort working with open-source communities through issues, pull request discussions, and code review.
  • BS/MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.

Preferred Qualifications

  • Open-source contributions to vLLM, SGLang, PyTorch, Triton, NCCL, Dynamo, or adjacent serving and runtime projects.
  • Experience delivering performance improvements such as attention or KV-cache efficiency, speculative decoding, scheduler improvements, quantization-aware serving, or streaming latency reductions.
  • Experience building reproducible benchmarking and performance regression infrastructure for latency and throughput.
  • Systems performance background spanning memory bandwidth, kernel fusion, PCIe/NVLink effects, and network fabrics such as InfiniBand.

Compensation And Benefits

The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Base salary is determined based on location, experience, and the pay of employees in similar positions. The position also includes eligibility for equity and benefits.

Applications will be accepted at least until July 20, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive, equal-opportunity work environment.

More jobs at Nvidia

Similar jobs