Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 4
Debugging @ 7
GPU
InfiniBand @ 4
LLM @ 4
Machine Learning
NVLink @ 4
Networking @ 8
Profiling @ 7
Reinforcement Learning @ 6
Rust @ 7
SGLang @ 4
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
An applied research team within NVIDIA’s Networking Systems & Software Architecture group is solving infrastructure problems for AI. The team builds systems-level software that moves data between GPUs, nodes, and storage, spanning low-level transport optimization, hardware-software co-design, and communication frameworks that integrate directly with production AI stacks. The team’s charter also includes emerging domains such as quantum computing interconnects.
The Senior Architect owns modules and projects end-to-end, from scoping research questions to shipping production code. The role requires a recognized expert who drives technical decisions, incorporates ideas from research and industry, and regularly prototypes new approaches to validate them. The work sits at the boundary of applied research and production engineering.
Responsibilities
- Architect and implement high-performance communication and memory management libraries for distributed AI.
- Drive hardware-software co-optimization with GPU, DPU, NIC, and switch teams using GPUDirect RDMA, NVLink, and next-generation interconnects.
- Profile and optimize data movement across GPU memory, system DRAM, NVMe, and network fabrics.
- Integrate networking capabilities into AI serving stacks such as vLLM, SGLang, and TensorRT-LLM.
- Contribute to and maintain open-source projects.
- Mentor engineers and conduct design reviews.
- Prototype experimental technologies and evaluate their viability.
Requirements
- 12+ years of experience in systems software and/or networking, with demonstrated ownership of complex projects.
- MS, PhD, or equivalent experience in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Solid understanding of high-performance networking, including InfiniBand, RoCE, RDMA, NVLink, and GPUDirect.
- Strong C, C++, and/or Rust systems programming skills, with comfort in performance profiling and low-level debugging.
- Understanding of ML systems concepts, including transformer architectures, KV cache mechanics, model parallelism, and distributed training or inference patterns.
Preferred Qualifications
- Knowledge of ML inference frameworks such as vLLM, SGLang, and TensorRT-LLM, including their communication requirements.
- Knowledge of storage networking, including NVMe-oF, GPUDirect Storage, and S3.
- Background in reinforcement learning systems.
Compensation and Benefits
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Base salary range for Level 6: USD 272,000–431,250 per year.
- Compensation is determined based on location, experience, and the pay of employees in similar positions.
- Eligible employees also receive equity and benefits.
- Applications will be accepted at least until May 23, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes.
- NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.