Senior Deep Learning Communication Architect

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI Algorithms CUDA @ 3 Deep Learning @ 7 GPU @ 3 InfiniBand @ 3 LLM @ 4 MPI NCCL NVLink OpenCL @ 3 PyTorch @ 6 Python @ 7 SGLang @ 6 TensorRT @ 6 vLLM @ 6

Details

NVIDIA's software architecture group is hiring a Deep Learning Communication Architect to scale deep neural network (DNN) models and training and inference frameworks to systems with hundreds of thousands of nodes.

Responsibilities

  • Optimize communication performance by identifying and eliminating bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
  • Design and implement communication algorithms and protocols for deep learning workloads, minimizing communication overhead and latency.
  • Collaborate with hardware and software teams to develop systems using high-speed interconnects such as NVLink, InfiniBand, and SPC-X, along with communication libraries including MPI, NCCL, UCX, UCC, and NVSHMEM.
  • Research and evaluate communication technologies and techniques to improve the performance and scalability of deep learning systems.
  • Build proofs of concept, conduct experiments, and perform quantitative modeling to validate and deploy new communication strategies.

Requirements

  • Ph.D., master's degree, bachelor's degree in Computer Science, Electrical Engineering, Computer Science and Electrical Engineering, or a closely related field, or equivalent experience.
  • At least 6 years of experience building and scaling DNNs, working with parallelism in DNN frameworks, or supporting deep learning training and inference workloads.
  • Experience evaluating, analyzing, and optimizing LLM training and inference performance for state-of-the-art models on cutting-edge hardware.
  • Deep understanding of data parallelism, pipeline parallelism, tensor parallelism, expert parallelism, and FSDP.
  • Understanding of emerging serving architectures such as disaggregated serving and inference servers including Dynamo and Triton.
  • Proficiency developing code for one or more DNN training and inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, or SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and with InfiniBand and RoCE networks.

Preferred Qualifications

  • Contributions to one or more DNN training and inference frameworks in previous work.
  • Deep understanding of, and contributions to, scaling LLMs on large-scale systems.

Compensation And Benefits

The base salary range is $184,000–$287,500 USD for Level 4 and $224,000–$356,500 USD for Level 5. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until May 24, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs