Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 4
Communication @ 6
Deep Learning @ 4
GPU
HPC @ 4
JAX @ 4
LLM @ 4
MPI @ 4
NCCL @ 6
Parallel Programming @ 4
Performance Analysis
PyTorch @ 4
Python @ 4
Reinforcement Learning @ 6
SGLang @ 4
System Architecture @ 7
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is leading groundbreaking developments in artificial intelligence, high-performance computing, and visualization. The GPU serves as the visual cortex of modern computers and is at the heart of NVIDIA's products and services.
This role focuses on bringing advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, and JAX. You will work with the team responsible for communication libraries such as NCCL and NVSHMEM, as well as technologies such as GPUDirect, supporting the scaling of deep learning and HPC applications. Customers have diverse multi-GPU requirements, ranging from training on clusters of up to 100,000 GPUs to inference workloads with microsecond latency.
Responsibilities
- Integrate new communication library features into AI frameworks, from proof of concept and performance analysis through production.
- Analyze AI workloads and frameworks to identify multi-GPU communication requirements and opportunities.
- Collaborate hands-on with teams working on the latest AI models.
- Improve AI compilers to hide communications or perform automatic fusion.
- Conduct in-depth performance characterization of AI workloads on multi-GPU clusters.
- Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
- Author custom communication or fused compute-communication kernels to achieve maximum performance on NVIDIA platforms.
- Influence the roadmap of communication libraries, including NCCL and NVSHMEM.
- Collaborate with a dynamic team across multiple time zones.
Requirements
- Bachelor's, master's, or doctoral degree in computer science or a related field, or equivalent experience, with more than five years of software engineering and HPC/AI experience.
- Development or integration experience with deep learning frameworks such as PyTorch and JAX, and inference engines such as TRT-LLM, vLLM, and SGLang.
- Rapid prototyping and development experience with Python, C++, CUDA, or related domain-specific languages such as Triton and cuTe.
- Strong understanding of AI models, parallelism, and/or compiler technologies such as
torch.compile. - Experience benchmarking performance on AI clusters.
- Familiarity with at least one performance profiler toolchain, such as PyTorch Profiler or NVIDIA Nsight Systems.
- Understanding of HPC/AI communication concepts, including one-sided versus two-sided communication, elasticity, resiliency, and topology discovery.
- Adaptability and enthusiasm for learning new areas and tools.
- Ability to work and communicate effectively across teams and time zones.
Preferred Qualifications
- Experience with parallel programming using at least one communication runtime, such as NCCL, NVSHMEM, or MPI.
- Strong understanding of computer system architecture, hardware-software interactions, operating system principles, and systems software fundamentals.
- Expertise in one or more of the following areas: training, distributed inference, mixture-of-experts, reinforcement learning, or kernel authoring with CUDA, Triton, cuTe, or similar technologies.
- Experience programming for compute and communication overlap in distributed runtimes.
- Experience with AI compiler pattern matching and lowering.
- Strong understanding of memory hierarchy, consistency models, and tensor layout.
Compensation and Benefits
The base salary depends on location, experience, and the pay of employees in similar positions. The base salary ranges are:
- Level 3: USD 152,000–241,500 per year
- Level 4: USD 184,000–287,500 per year
The role is also eligible for equity and benefits. Applications will be accepted at least until August 3, 2026. NVIDIA is an equal opportunity employer and uses AI tools in its recruiting processes.