Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
CUDA @ 6
Communication @ 7
GPU @ 7
Networking @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior GPU Networking Architect to join its networking software group, bringing strong GPU architecture and programming skills to build and improve GPU communication kernels. This role links GPU computing with networking by developing communication primitives alongside GPU hardware capabilities.
Responsibilities
- Build, implement, and optimize GPU communication kernels that underpin collective and point-to-point operations in large-scale AI systems.
- Leverage deep knowledge of GPU architecture—thread scheduling, memory hierarchy, execution pipelines—to improve kernel efficiency, minimize latency, and overlap computation with communication.
- Develop GPU-resident communication primitives and device-side APIs that enable fine-grained, kernel-initiated data movement across nodes and accelerators.
- Profile and tune GPU kernels end-to-end, identifying bottlenecks at the intersection of compute, memory, and network, and driving targeted optimizations.
- Collaborate with network software, hardware, and AI framework teams to co-design communication strategies that align with GPU execution patterns and emerging model architectures.
- Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and validate new communication strategies before committing them to production.
- Contribute to the evolution of programming models that expose GPU-aware networking capabilities to application developers.
Requirements
- 5+ years of hands-on CUDA programming, including writing and optimizing non-trivial GPU kernels.
- M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field.
- Strong understanding of GPU architecture fundamentals: warp scheduling, shared memory, L2 cache, memory coalescing, occupancy tuning, and asynchronous execution.
- Experience with systems-level C/C++ development in performance-critical environments.
- Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication.
- Ability to read and reason about GPU performance profiles (e.g., Nsight Compute, Nsight Systems) and translate observations into actionable optimizations.
- Strong collaboration skills in a multi-national, interdisciplinary environment.
Benefits
- NVIDIA offers highly competitive salaries and a comprehensive benefits package (see www.nvidiabenefits.com/).
More jobs at Nvidia
Senior System Software Engineer, Interactive World Models
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer – Streaming
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Software Engineer - GeForce Now Low Latency Streaming
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Gpu Architecture Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Compiler Engineer, Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year
Senior HPC And AI Network Software Architect
Nvidia · Zurich, Switzerland
PLN 221,200-507,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Marketing Engineer, Enterprise AI Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Technical Support Engineer, Tavily
Nebius · Tel Aviv, Israel, World
USD 109,500-136,800 per year
Head of Forward Deployment Engineering - Tavily
Nebius · United States, New York City, United States
USD 230,000-310,000 per year