Principal Deep Learning Communication Architect

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 Agentic AI @ 6 Algorithms CUDA @ 4 Communication @ 4 Deep Learning @ 8 GPU @ 4 HPC InfiniBand @ 4 JAX @ 4 LLM @ 7 MPI @ 6 NCCL @ 6 NVLink Networking @ 6 PyTorch @ 6 SGLang @ 7 Technical Proficiency @ 6 TensorRT @ 7 vLLM @ 7

Details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI to define the next era of computing, with GPUs powering computers, robots, and self-driving cars.

Responsibilities

  • Define the long-term technical roadmap for communication libraries across NVIDIA’s next-generation platforms, ensuring seamless scaling of models to clusters comprising hundreds of thousands of nodes.
  • Lead the development of next-generation communication primitives and collective algorithms optimized for heterogeneous interconnects, including NVLink, Spectrum-X Ethernet, and Quantum-X InfiniBand.
  • Partner with application developers to architect and implement specialized communication primitives for AI and HPC libraries, including NCCL, NIXL, NVSHMEM, UCC, and UCX, supporting trillion-parameter models and Agentic AI.
  • Collaborate with silicon architects and software engineers to influence hardware specifications for next-generation networking and meet the demands of trillion-parameter large language models and Agentic AI.
  • Develop high-fidelity analytical models and simulators to predict system behavior under emerging workloads.

Requirements

  • Ph.D. or M.S. in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • More than 12 years of industry experience in high-performance computing or distributed deep learning.
  • Deep understanding of 3D parallelism, including data, tensor, and pipeline parallelism, as well as context parallelism, expert parallelism, and Zero Redundancy Optimizer variants.
  • Deep technical proficiency with NCCL, UCX, UCC, NVSHMEM, or MPI.
  • Experience with RDMA, RoCE, and low-level InfiniBand verbs.
  • Advanced knowledge of high-throughput inference engines and schedulers, specifically TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
  • Expert knowledge of the NVIDIA GPU memory hierarchy, including HBM3e, HBM4, and L2 cache, as well as CUDA programming models.

Preferred Qualifications

  • Hands-on experience developing with Megatron-Core, DeepSpeed, or JAX/XLA, including an understanding of how these frameworks interact with low-level communication runtimes.
  • Significant upstream contributions to major open-source projects such as PyTorch Distributed, KServe, or Ray.
  • A track record of deploying and optimizing models on NVIDIA platforms or similar rack-scale systems.
  • Patents or papers in top-tier systems and architecture venues such as ISCA, ASPLOS, NeurIPS, or SC.

Benefits

The role offers competitive salaries, equity, and a generous benefits package. NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer. NVIDIA also uses AI tools in its recruiting processes.

Applications will be accepted at least until April 18, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs