Senior GPU Networking Architect

at Nvidia
PLN 292,500-650,000 per year
SENIOR
✅ On-site

Tech Stack

AI API CUDA @ 6 Communication @ 7 Deep Learning @ 4 GPU @ 7 InfiniBand @ 6 LLM @ 4 NCCL @ 4 NVLink @ 6 Networking @ 6 PyTorch @ 4 TensorRT @ 4 vLLM @ 4

Details

NVIDIA is developing the software foundation for large-scale AI systems and is seeking a Senior GPU Networking Architect to join its networking software group. The role combines GPU computing and networking by developing communication primitives alongside GPU hardware capabilities.

Responsibilities

  • Build, implement, and optimize GPU communication kernels supporting collective and point-to-point operations in large-scale AI systems.
  • Use deep knowledge of GPU architecture, including thread scheduling, memory hierarchy, and execution pipelines, to improve kernel efficiency, minimize latency, and overlap computation with communication.
  • Develop GPU-resident communication primitives and device-side APIs for fine-grained, kernel-initiated data movement across nodes and accelerators.
  • Profile and tune GPU kernels end to end, identifying bottlenecks across compute, memory, and networking and implementing targeted optimizations.
  • Collaborate with network software, hardware, and AI framework teams to co-design communication strategies aligned with GPU execution patterns and emerging model architectures.
  • Build proofs of concept, conduct experiments, and perform quantitative modeling to evaluate new communication strategies before production implementation.
  • Contribute to programming models that expose GPU-aware networking capabilities to application developers.

Requirements

  • 5 or more years of hands-on CUDA programming, including writing and optimizing non-trivial GPU kernels.
  • M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field.
  • Strong understanding of GPU architecture fundamentals, including warp scheduling, shared memory, L2 cache, memory coalescing, occupancy tuning, and asynchronous execution.
  • Experience with systems-level C/C++ development in performance-critical environments.
  • Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication.
  • Ability to analyze GPU performance profiles using tools such as Nsight Compute and Nsight Systems and translate findings into actionable optimizations.
  • Strong collaboration skills in a multinational, interdisciplinary environment.

Preferred Qualifications

  • Experience developing or optimizing communication kernels in NCCL, NVSHMEM, or similar GPU-aware communication frameworks.
  • Understanding of distributed deep learning parallelism techniques, including data, tensor, pipeline, expert, and mixture-of-experts parallelism, as well as their communication patterns.
  • Background in RDMA, InfiniBand, high-speed networking, and GPU system topology, including NVLink, NVSwitch, PCIe, and network fabrics.
  • Experience with kernel pipelining, persistent kernels, or cooperative groups to hide communication latency behind computation.
  • Experience evaluating and optimizing large-scale LLM training or inference workloads using frameworks such as PyTorch, TensorRT-LLM, or vLLM.
  • Familiarity with emerging serving architectures such as disaggregated serving.

Benefits

NVIDIA offers competitive salaries and a comprehensive benefits package for employees and their families. More information is available at www.nvidiabenefits.com.

Compensation

For Poland, the base salary range is 292,500 PLN–507,000 PLN for Level 4 and 375,000 PLN–650,000 PLN for Level 5. The base salary is determined by location, experience, and compensation for employees in similar positions.

More jobs at Nvidia

Similar jobs