Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms
CUDA @ 3
Deep Learning @ 7
GPU @ 3
InfiniBand @ 3
LLM @ 4
MPI
NCCL
NVLink
OpenCL @ 3
PyTorch @ 6
Python @ 7
SGLang @ 6
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's software architecture group is hiring a Deep Learning Communication Architect to scale deep neural network (DNN) models and training and inference frameworks to systems with hundreds of thousands of nodes.
Responsibilities
- Optimize communication performance by identifying and eliminating bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
- Design and implement communication algorithms and protocols for deep learning workloads, minimizing communication overhead and latency.
- Collaborate with hardware and software teams to develop systems using high-speed interconnects such as NVLink, InfiniBand, and SPC-X, along with communication libraries including MPI, NCCL, UCX, UCC, and NVSHMEM.
- Research and evaluate communication technologies and techniques to improve the performance and scalability of deep learning systems.
- Build proofs of concept, conduct experiments, and perform quantitative modeling to validate and deploy new communication strategies.
Requirements
- Ph.D., master's degree, bachelor's degree in Computer Science, Electrical Engineering, Computer Science and Electrical Engineering, or a closely related field, or equivalent experience.
- At least 6 years of experience building and scaling DNNs, working with parallelism in DNN frameworks, or supporting deep learning training and inference workloads.
- Experience evaluating, analyzing, and optimizing LLM training and inference performance for state-of-the-art models on cutting-edge hardware.
- Deep understanding of data parallelism, pipeline parallelism, tensor parallelism, expert parallelism, and FSDP.
- Understanding of emerging serving architectures such as disaggregated serving and inference servers including Dynamo and Triton.
- Proficiency developing code for one or more DNN training and inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, or SGLang.
- Strong programming skills in C++ and Python.
- Familiarity with GPU computing, including CUDA and OpenCL, and with InfiniBand and RoCE networks.
Preferred Qualifications
- Contributions to one or more DNN training and inference frameworks in previous work.
- Deep understanding of, and contributions to, scaling LLMs on large-scale systems.
Compensation And Benefits
The base salary range is $184,000–$287,500 USD for Level 4 and $224,000–$356,500 USD for Level 5. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until May 24, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Principal Deep Learning Communication Architect
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Deep Learning Frameworks CUDA Software Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Architect, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year