Senior Scale-Up Network System Architect

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 6 GPU HPC @ 6 LLM NVLink @ 4 Networking @ 4 System Architecture @ 4

Details

NVIDIA is seeking a Senior Scale-Up Network System Architect to develop the next generation of NVLink-based scale-up networking for leading AI supercomputing platforms. The role focuses on the complete architecture of the interconnect linking GPUs within and between racks, enabling large AI training and inference clusters to operate as a single accelerator. This system-level position involves collaboration across silicon, firmware, software, and topology development teams.

Responsibilities

  • Define end-to-end system architecture for next-generation NVLink scale-up networks, from link and switch behavior through rack- and pod-level topology.
  • Convert AI training and inference workload requirements, including LLM, MoE, and emerging model architectures, into networking, bandwidth, latency, and resiliency specifications.
  • Drive architecture trade-off studies across performance, cost, power, and reliability, and build data-driven models to support architectural decisions.
  • Collaborate with ASIC, firmware, software, and systems teams to ensure architecture decisions are implementable and accurately delivered in silicon and products.
  • Develop and use simulation and analytical models to validate architecture choices before committing to silicon.
  • Represent system architecture in multifunctional and customer-facing technical discussions.
  • Serve as a technical expert on scale-up network behavior at scale.

Requirements

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • 8 or more years of industry experience in computer networking, system architecture, or high-performance interconnects.
  • Strong understanding of network structure and dynamics at scale, including the effects of topology, congestion, and failure modes on distributed workload performance.
  • Experience developing or using simulation and modeling environments to evaluate architectural trade-offs.
  • Strong multifunctional collaboration skills, with the ability to drive alignment across hardware, firmware, and software teams without direct authority.
  • Clear technical communication skills, including the ability to explain complex architectural trade-offs to engineering and non-engineering collaborators.

Preferred Qualifications

  • Direct experience with NVLink, NVSwitch, or comparable scale-up interconnect technologies.
  • Hands-on experience with high-performance networking transports such as RoCE (RDMA over Converged Ethernet), including low-latency transport, congestion control, and lossless fabric design.
  • Experience with memory subsystem architecture, including memory hierarchies, cache coherency, and HBM, DDR, or LPDDR memory technologies.
  • Working knowledge of large-scale AI model architectures and how they map to network requirements.
  • Background in HPC or supercomputing-scale interconnect development.

Compensation and Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until October 8, 2026. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs