Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 6
GPU
HPC @ 6
LLM
NVLink @ 4
Networking @ 4
System Architecture @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Scale-Up Network System Architect to develop the next generation of NVLink-based scale-up networking for leading AI supercomputing platforms. The role focuses on the complete architecture of the interconnect linking GPUs within and between racks, enabling large AI training and inference clusters to operate as a single accelerator. This system-level position involves collaboration across silicon, firmware, software, and topology development teams.
Responsibilities
- Define end-to-end system architecture for next-generation NVLink scale-up networks, from link and switch behavior through rack- and pod-level topology.
- Convert AI training and inference workload requirements, including LLM, MoE, and emerging model architectures, into networking, bandwidth, latency, and resiliency specifications.
- Drive architecture trade-off studies across performance, cost, power, and reliability, and build data-driven models to support architectural decisions.
- Collaborate with ASIC, firmware, software, and systems teams to ensure architecture decisions are implementable and accurately delivered in silicon and products.
- Develop and use simulation and analytical models to validate architecture choices before committing to silicon.
- Represent system architecture in multifunctional and customer-facing technical discussions.
- Serve as a technical expert on scale-up network behavior at scale.
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8 or more years of industry experience in computer networking, system architecture, or high-performance interconnects.
- Strong understanding of network structure and dynamics at scale, including the effects of topology, congestion, and failure modes on distributed workload performance.
- Experience developing or using simulation and modeling environments to evaluate architectural trade-offs.
- Strong multifunctional collaboration skills, with the ability to drive alignment across hardware, firmware, and software teams without direct authority.
- Clear technical communication skills, including the ability to explain complex architectural trade-offs to engineering and non-engineering collaborators.
Preferred Qualifications
- Direct experience with NVLink, NVSwitch, or comparable scale-up interconnect technologies.
- Hands-on experience with high-performance networking transports such as RoCE (RDMA over Converged Ethernet), including low-latency transport, congestion control, and lossless fabric design.
- Experience with memory subsystem architecture, including memory hierarchies, cache coherency, and HBM, DDR, or LPDDR memory technologies.
- Working knowledge of large-scale AI model architectures and how they map to network requirements.
- Background in HPC or supercomputing-scale interconnect development.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until October 8, 2026. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.