Senior Deep Learning Framework Communications Engineer

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 4 Communication @ 6 Deep Learning @ 4 GPU HPC @ 4 JAX @ 4 LLM @ 4 MPI @ 4 NCCL @ 6 Parallel Programming @ 4 Performance Analysis PyTorch @ 4 Python @ 4 Reinforcement Learning @ 6 SGLang @ 4 System Architecture @ 7 vLLM @ 4

Details

NVIDIA is leading groundbreaking developments in artificial intelligence, high-performance computing, and visualization. The GPU serves as the visual cortex of modern computers and is at the heart of NVIDIA's products and services.

This role focuses on bringing advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, and JAX. You will work with the team responsible for communication libraries such as NCCL and NVSHMEM, as well as technologies such as GPUDirect, supporting the scaling of deep learning and HPC applications. Customers have diverse multi-GPU requirements, ranging from training on clusters of up to 100,000 GPUs to inference workloads with microsecond latency.

Responsibilities

  • Integrate new communication library features into AI frameworks, from proof of concept and performance analysis through production.
  • Analyze AI workloads and frameworks to identify multi-GPU communication requirements and opportunities.
  • Collaborate hands-on with teams working on the latest AI models.
  • Improve AI compilers to hide communications or perform automatic fusion.
  • Conduct in-depth performance characterization of AI workloads on multi-GPU clusters.
  • Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
  • Author custom communication or fused compute-communication kernels to achieve maximum performance on NVIDIA platforms.
  • Influence the roadmap of communication libraries, including NCCL and NVSHMEM.
  • Collaborate with a dynamic team across multiple time zones.

Requirements

  • Bachelor's, master's, or doctoral degree in computer science or a related field, or equivalent experience, with more than five years of software engineering and HPC/AI experience.
  • Development or integration experience with deep learning frameworks such as PyTorch and JAX, and inference engines such as TRT-LLM, vLLM, and SGLang.
  • Rapid prototyping and development experience with Python, C++, CUDA, or related domain-specific languages such as Triton and cuTe.
  • Strong understanding of AI models, parallelism, and/or compiler technologies such as torch.compile.
  • Experience benchmarking performance on AI clusters.
  • Familiarity with at least one performance profiler toolchain, such as PyTorch Profiler or NVIDIA Nsight Systems.
  • Understanding of HPC/AI communication concepts, including one-sided versus two-sided communication, elasticity, resiliency, and topology discovery.
  • Adaptability and enthusiasm for learning new areas and tools.
  • Ability to work and communicate effectively across teams and time zones.

Preferred Qualifications

  • Experience with parallel programming using at least one communication runtime, such as NCCL, NVSHMEM, or MPI.
  • Strong understanding of computer system architecture, hardware-software interactions, operating system principles, and systems software fundamentals.
  • Expertise in one or more of the following areas: training, distributed inference, mixture-of-experts, reinforcement learning, or kernel authoring with CUDA, Triton, cuTe, or similar technologies.
  • Experience programming for compute and communication overlap in distributed runtimes.
  • Experience with AI compiler pattern matching and lowering.
  • Strong understanding of memory hierarchy, consistency models, and tensor layout.

Compensation and Benefits

The base salary depends on location, experience, and the pay of employees in similar positions. The base salary ranges are:

  • Level 3: USD 152,000–241,500 per year
  • Level 4: USD 184,000–287,500 per year

The role is also eligible for equity and benefits. Applications will be accepted at least until August 3, 2026. NVIDIA is an equal opportunity employer and uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs