Principal Software Engineer – Large-Scale LLM Memory and Storage Systems

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 3 Communication @ 4 Distributed Systems @ 8 GPU @ 7 GenAI Generative AI LLM @ 6 Machine Learning NVLink @ 3 Networking Profiling @ 7 Python @ 8 Rust SGLang TensorRT vLLM

Details

NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads.

The team is seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems.

Responsibilities

  • Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file, object, and cloud storage to support large-scale LLM inference.
  • Architect and implement deep integrations with leading LLM serving engines such as vLLM, SGLang, and TensorRT-LLM, focusing on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.
  • Co-design interfaces and protocols for disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage across GPU, CPU, local disk, and remote memory for high-throughput, low-latency inference.
  • Partner with GPU architecture, networking, and platform teams to use GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.
  • Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums, including open source, conferences, and customer-facing technical deep dives.

Requirements

  • Master's degree, PhD, or equivalent experience.
  • 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services.
  • Deep understanding of memory hierarchies, including GPU HBM, host DRAM, SSD, and remote or object storage, along with experience designing systems spanning multiple tiers for performance and cost efficiency.
  • Experience with distributed caching or key-value systems, especially designs optimized for low latency and high concurrency.
  • Hands-on experience with networked I/O and RDMA, NVMe-oF, or NVLink-style technologies, and familiarity with disaggregated and aggregated deployments for AI clusters.
  • Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in time to first token (TTFT) and throughput.
  • Excellent communication skills and experience leading cross-functional efforts with research, product, and customer teams.

Preferred Qualifications

  • Contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse.
  • Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, particularly in enterprise or hyperscale environments.
  • Publications or patents in LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML.

Benefits

  • Competitive salary and comprehensive benefits package.
  • Eligibility for equity and NVIDIA benefits.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

Applications for this job will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs