Senior Software Architect - Deep Learning and HPC Communications
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms @ 4
CUDA @ 4
Communication @ 4
Debugging @ 6
Deep Learning @ 4
GPU
HPC @ 4
InfiniBand @ 4
Linux @ 7
MPI @ 4
NCCL @ 4
NVLink @ 4
Networking
Parallel Programming @ 4
PyTorch @ 4
System Architecture @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is advancing developments in artificial intelligence, high-performance computing, and visualization. The GPU Communications Libraries and Networking team builds communication libraries such as NCCL, NVSHMEM, and UCX, which are essential for scaling deep learning and HPC workloads. The team is seeking a Senior Software Architect to help co-design next-generation data center platforms and scalable communications software.
Deep learning and HPC applications have substantial compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected through high-speed interconnects such as NVLink and PCIe within a node, and through high-speed networking such as InfiniBand and Ethernet across nodes. Efficient communication between GPUs directly affects end-to-end application performance, particularly as systems continue to scale.
Responsibilities
- Investigate opportunities to improve communication performance by identifying bottlenecks in current systems.
- Design and implement new communication technologies to accelerate AI and HPC workloads.
- Explore innovative hardware and software solutions for next-generation platforms through co-design efforts involving GPU, networking, and software architects.
- Build proofs of concept, conduct experiments, and perform quantitative modeling to evaluate and drive new innovations.
- Use simulation to explore the performance of large GPU clusters at scales of hundreds of thousands of GPUs.
Requirements
- M.S. or Ph.D. degree in Computer Science, Computer Engineering, or equivalent experience.
- Five or more years of relevant experience.
- Excellent C and C++ programming and debugging skills.
- Experience with parallel programming models such as MPI and SHMEM, and at least one communication runtime, including MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, or UCC.
- Deep understanding of operating systems, computer architecture, and system architecture.
- Solid understanding of network architecture, topology, algorithms, and communication scaling relevant to AI and HPC workloads.
- Strong experience with Linux.
- Ability and flexibility to work and communicate effectively in a multinational, multi-time-zone corporate environment.
Preferred Qualifications
- Expertise in related technologies and a passion for the work, including experience with CUDA programming and NVIDIA GPUs.
- Knowledge of high-performance networks such as InfiniBand, RoCE, and NVLink.
- Experience with deep learning frameworks such as PyTorch and TensorFlow.
- Knowledge of deep learning parallelism and its mapping to the communication subsystem.
- Experience with HPC applications.
- Strong collaborative and interpersonal skills, with a proven ability to guide and influence others in a dynamic, multifunctional environment.
Compensation and Benefits
- Base salary range for Level 4: USD 184,000–287,500 per year.
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until August 1, 2026.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.