Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 7
Communication @ 4
GPU @ 7
HPC @ 6
HTTP
JAX @ 4
LLM
NCCL @ 6
Networking @ 6
PyTorch @ 4
Python @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior HPC and AI Network Software Architect to help build the next generation of scalable AI infrastructure. The role focuses on distributed training, real-time inference, and communication efficiency across large systems. You will develop software and hardware approaches, shape platform evolution through hands-on innovation, and contribute to systems powering high-performance AI workloads at scale. You will collaborate with researchers and engineers building software and hardware for AI infrastructure.
Responsibilities
- Build and evolve the architecture of scalable software systems for distributed AI training and inference, focusing on throughput, latency, resiliency, and memory efficiency across cluster-scale deployments.
- Develop and evaluate next-generation communication and runtime capabilities in libraries such as NCCL, UCX, and UCC for frontier AI workloads.
- Partner with AI framework teams, including TensorFlow, PyTorch, and JAX, as well as internal platform teams to build integrations, explore new approaches, and improve end-to-end performance and reliability.
- Collaborate on hardware and system-level features across GPUs, DPUs, and interconnects to accelerate data movement and enable training, inference, and model serving at scale.
- Drive innovation across runtime systems, communication libraries, and AI-specific protocol layers, turning new ideas into practical capabilities and robust implementations.
Requirements
- Ph.D. or equivalent industry experience in computer science, computer engineering, or a closely related field.
- At least 5 years of experience in systems programming, parallel or distributed computing, high-performance networking, or large-scale data movement, including experience designing and building complex systems.
- Strong programming skills in C++ and Python, with ideally CUDA or other GPU programming model experience, and a track record of building production-quality, performance-critical software.
- Extensive hands-on experience with AI frameworks such as PyTorch, TensorFlow, or JAX, along with an understanding of how communication libraries and runtime systems support large-scale training and inference.
- Demonstrated success developing and refining high-throughput, low-latency systems, with the ability to reason across software stacks, hardware capabilities, and system bottlenecks.
- Strong collaboration skills in a multinational, interdisciplinary environment, including the ability to work effectively with senior engineers, researchers, and partner teams.
Preferred Qualifications
- Deep expertise with NCCL, UCX, UCC, or similar communication libraries used in large-scale AI and HPC workloads.
- Strong background in networking and communication protocols, RDMA, collective communications, congestion-aware transport, or accelerator-aware networking.
- Comprehensive knowledge of large-model training and inference serving at scale, including communication bottlenecks, scheduling challenges, and system-level tradeoffs across compute, memory, and fabric.
- Experience with hardware-software co-design for distributed AI systems, including contributions to GPU, DPU, interconnect, or runtime capabilities.
- Familiarity with infrastructure for deploying LLMs or transformer-based models, including sharding, pipelining, expert parallelism, or hybrid parallelism.
Benefits
NVIDIA offers highly competitive salaries and a comprehensive benefits package. Additional information is available at www.nvidiabenefits.com.
Compensation
For Poland, the base salary range is PLN 221,250–383,500 for Level 3 and PLN 292,500–507,000 for Level 4. The base salary is determined based on location, experience, and the pay of employees in similar positions.