Senior Software Architect - Deep Learning and HPC Communications

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 Algorithms @ 6 CUDA @ 4 Communication @ 4 Debugging @ 6 Deep Learning @ 4 GPU HPC @ 4 InfiniBand @ 4 Linux @ 7 MPI @ 4 NCCL @ 4 NVLink @ 4 Networking Parallel Programming @ 4 PyTorch @ 4 System Architecture @ 7 TensorFlow @ 4

Details

NVIDIA's GPU Communications Libraries and Networking team builds communication libraries such as NCCL, NVSHMEM, and UCX, which are crucial for scaling deep learning and high-performance computing workloads. The team is seeking a Senior Software Architect to help co-design next-generation data center platforms and scalable communications software.

Deep learning and HPC applications have significant compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects such as NVLink and PCIe within a node, and with high-speed networking such as InfiniBand and Ethernet across nodes. Efficient communication between GPUs directly impacts end-to-end application performance, and this impact continues to grow as systems scale.

Responsibilities

  • Investigate opportunities to improve communication performance by identifying bottlenecks in current systems.
  • Design and implement new communication technologies to accelerate AI and HPC workloads.
  • Explore innovative hardware and software solutions for next-generation platforms as part of co-design efforts involving GPU, networking, and software architects.
  • Build proofs of concept, conduct experiments, and perform quantitative modeling to evaluate and drive new innovations.
  • Use simulation to explore the performance of large GPU clusters at scales of hundreds of thousands of GPUs.

Requirements

  • M.S. or Ph.D. degree in Computer Science, Computer Engineering, or equivalent experience.
  • 12 or more years of relevant experience.
  • Excellent C and C++ programming and debugging skills.
  • Experience with parallel programming models such as MPI and SHMEM, and at least one communication runtime, including MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, or UCC.
  • Deep understanding of operating systems, computer architecture, and system architecture.
  • Solid fundamentals in network architecture, topology, algorithms, and communication scaling relevant to AI and HPC workloads.
  • Strong experience with Linux.
  • Ability and flexibility to work and communicate effectively in a multinational, multi-time-zone corporate environment.

Preferred Qualifications

  • Expertise in related technologies and passion for the work.
  • Experience with CUDA programming and NVIDIA GPUs.
  • Knowledge of high-performance networks such as InfiniBand, RoCE, and NVLink.
  • Experience with deep learning frameworks such as PyTorch and TensorFlow.
  • Knowledge of deep learning parallelism and its mapping to the communication subsystem.
  • Experience with HPC applications.
  • Strong collaborative and interpersonal skills, with a proven track record of guiding and influencing others in a dynamic, multifunctional environment.

Compensation and Benefits

The base salary range is USD 224,000–356,500 for Level 5 and USD 272,000–431,250 for Level 6. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until July 16, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs