Distinguished Software Architect - Deep Learning and HPC Communications

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI Algorithms @ 6 CUDA @ 6 Communication @ 6 Debugging @ 7 Deep Learning @ 6 GPU @ 6 HPC @ 6 InfiniBand @ 7 Leadership @ 6 MPI @ 6 Machine Learning @ 6 NCCL @ 6 NVLink Networking @ 7 Parallel Programming @ 6 Performance Analysis @ 7 PyTorch @ 6 Software Development @ 6 System Architecture @ 6 TensorFlow @ 6

Details

NVIDIA is leading groundbreaking developments in artificial intelligence, high-performance computing, and visualization. The GPU, NVIDIA's invention, serves as the visual cortex of modern computers and is at the heart of its products and services.

The GPU Communications Libraries and Networking team delivers communication libraries such as NCCL, NVSHMEM, and UCX for deep learning and high-performance computing. The team is seeking a Distinguished Software Architect to help co-design next-generation data center platforms. Deep learning and HPC applications run at scales of up to tens of thousands of GPUs. GPUs are connected through high-speed interconnects such as NVLink and PCIe within a node, and through high-speed networking such as InfiniBand and Ethernet across nodes. Communication performance between GPUs directly affects end-to-end application performance, particularly at large scales.

Responsibilities

  • Research new communication technologies, including expanding the GPUDirect technology portfolio, and design new features for communication libraries.
  • Propose innovative hardware and software solutions for next-generation platforms.
  • Co-design solutions with GPU, networking, and software architects and ensure seamless integration with software stacks.
  • Inspire changes based on quantitative data from proof-of-concepts, detailed technical analysis, and modeling.
  • Drive adoption of new communication technologies across application verticals.
  • Keep up with the latest deep learning research.
  • Collaborate with diverse internal and external teams, including deep learning researchers and customers.

Requirements

  • PhD in Computer Science, Computer Engineering, or a related field, or strong equivalent experience.
  • 15+ years of relevant academic or industry experience.
  • Expertise in HPC, parallel programming models such as MPI and SHMEM, at least one communication runtime such as MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, or UCC, computer and system architecture, GPU architecture, and CUDA.
  • Deep understanding of high-performance networking, including InfiniBand, Ethernet, network design, network topologies, network debugging, and performance analysis.
  • Strength in several of the following areas: machine learning and deep learning fundamentals and their relationship to communications; parallel algorithms; fault tolerance and resiliency; competitive assessments; performance analysis and optimization for parallel applications on large clusters; and application development using deep learning frameworks such as PyTorch and TensorFlow.
  • Programming fluency in C or C++ for systems software development.
  • Ability to work and communicate effectively across hardware and software teams and time zones.

Preferred Qualifications

  • Industry-recognized leadership in HPC or deep learning communications, with a history of patents, publications, conference talks, or keynotes relevant to the role.
  • Influential contributions to industry standards such as MPI or OpenSHMEM and open-source software such as PyTorch, UCX, or Open MPI.

Benefits

  • Base salary range of $320,000–$488,750 USD per year, determined by location, experience, and the pay of employees in similar positions.
  • Eligibility for equity and benefits.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs