Senior GPU Networking Architect

at Nvidia
PLN 292,500-650,000 per year
SENIOR
✅ On-site

Tech Stack

AI API CUDA @ 6 Communication @ 7 GPU @ 7 Networking @ 7

Details

NVIDIA is looking for a Senior GPU Networking Architect to join its networking software group, bringing strong GPU architecture and programming skills to build and improve GPU communication kernels. This role links GPU computing with networking by developing communication primitives alongside GPU hardware capabilities.

Responsibilities

  • Build, implement, and optimize GPU communication kernels that underpin collective and point-to-point operations in large-scale AI systems.
  • Leverage deep knowledge of GPU architecture—thread scheduling, memory hierarchy, execution pipelines—to improve kernel efficiency, minimize latency, and overlap computation with communication.
  • Develop GPU-resident communication primitives and device-side APIs that enable fine-grained, kernel-initiated data movement across nodes and accelerators.
  • Profile and tune GPU kernels end-to-end, identifying bottlenecks at the intersection of compute, memory, and network, and driving targeted optimizations.
  • Collaborate with network software, hardware, and AI framework teams to co-design communication strategies that align with GPU execution patterns and emerging model architectures.
  • Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and validate new communication strategies before committing them to production.
  • Contribute to the evolution of programming models that expose GPU-aware networking capabilities to application developers.

Requirements

  • 5+ years of hands-on CUDA programming, including writing and optimizing non-trivial GPU kernels.
  • M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field.
  • Strong understanding of GPU architecture fundamentals: warp scheduling, shared memory, L2 cache, memory coalescing, occupancy tuning, and asynchronous execution.
  • Experience with systems-level C/C++ development in performance-critical environments.
  • Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication.
  • Ability to read and reason about GPU performance profiles (e.g., Nsight Compute, Nsight Systems) and translate observations into actionable optimizations.
  • Strong collaboration skills in a multi-national, interdisciplinary environment.

Benefits

More jobs at Nvidia

Similar jobs