Principal Architect, AI Networking

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 6 Communication @ 4 GPU @ 6 InfiniBand @ 8 LLM @ 4 MPI @ 4 Machine Learning @ 4 NCCL @ 4 NVLink @ 8 Networking @ 8 Python @ 6 Rust @ 6 SGLang @ 4 TensorRT @ 4 vLLM @ 4

Details

An applied research team within NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure problems. The team builds systems-level software that moves data between GPUs, nodes, and storage at the speed modern AI demands, spanning low-level transport optimization, hardware-software co-design, and communication frameworks that plug directly into production AI stacks. The team’s charter expands into emerging domains including quantum computing interconnects.

This Principal Architect role leads the research agenda and architectural direction for how NVIDIA’s AI systems communicate at scale across GPUs, DPUs, NICs, and heterogeneous storage. It requires someone who defines project scope from scratch, publishes original work, and translates research breakthroughs into production-grade software that ships industry-wide.

Responsibilities

  • Set the long-term technical vision for distributed AI communication systems, including GPU-to-GPU, GPU-to-storage, and cross-node data movement.
  • Conduct original research and prototype next-generation networking solutions over RDMA, NVLink, and GPUDirect.
  • Drive hardware-software co-optimization with GPUs, DPUs, NICs, and network switches.
  • Investigate fundamental bottlenecks in communication runtimes for large-scale AI workloads, including KV cache transfer, disaggregated prefill/decode, and model parallelism.
  • Integrate networking capabilities into AI serving stacks such as vLLM, SGLang, and TensorRT-LLM.
  • Publish findings and represent NVIDIA in industry forums and standards bodies.
  • Mentor senior engineers across the organization.

Requirements

  • 15+ years of experience in systems software and/or networking, with deep expertise in high-performance networking such as InfiniBand, RoCE, RDMA, and NVLink.
  • Experience with communication libraries including NIXL, NCCL, UCX, MPI, and NVSHMEM.
  • Expertise in GPU-accelerated systems and a track record of defining and delivering complex, cross-team technical initiatives from research concept to production.
  • MS, PhD, or equivalent experience in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Deep understanding of computer architecture, memory hierarchies, DMA engines, and OS-level networking.
  • Understanding of machine learning systems concepts, including transformer architectures, KV cache mechanics, model parallelism, and distributed training and inference patterns.
  • Proficiency in programming languages such as C, C++, Rust, and Python.

Preferred Qualifications

  • Knowledge of ML inference frameworks such as vLLM, SGLang, and TensorRT-LLM and their communication requirements.
  • CUDA programming and NVIDIA GPU architecture expertise.
  • Proven experience influencing product strategy and technical roadmaps at a senior level.
  • Major open-source contributions.

Benefits

NVIDIA offers competitive salaries, a comprehensive benefits package, equity, and benefits. NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.

The base salary range is 272,000 USD–431,250 USD. Applications will be accepted at least until April 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs