Principal Developer, AI Networking

at Nvidia
USD 272,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Bash @ 6 CUDA @ 4 Communication @ 6 Debugging Deep Learning @ 4 Distributed Systems GPU LLM @ 4 MPI @ 8 Machine Learning NCCL @ 8 Networking @ 8 Performance Analysis Profiling PyTorch @ 4 Python @ 6 TensorFlow @ 4

Details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing.

The AI Networking Codesign and Benchmarking R&D group requires a senior software engineer. In this exciting role, you will profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters used for distributed Deep Learning LLM training and inference. Your primary focus will be collectives communication and networking.

You will work across hardware components such as HCAs, Switches, CPUs, GPUs, and Systems. You will also engage with software layers including LLM applications, machine learning frameworks, communication, and computing libraries. Moreover, you will build performance analysis tools and strategies to investigate details and clarify performance expectations, limitations, and bottlenecks.

Responsibilities

  • Characterizing AI workloads and deep learning models aimed at large-scale LLM training and inference on NVIDIA supercomputers, with a focus on distributed systems and high-performance networking.
  • Benchmarking, profiling, and analyzing performance to find bottlenecks and identify areas for improvement and optimizations, emphasizing networking.
  • Developing PyTorch trace-based profiling, analysis, and replaying toolset to aid in benchmarking, debugging, and co-designing network systems for LLM workloads.
  • Collaborating with multiple teams from hardware to software to provide performance analysis insights.
  • Defining performance test plans, setting performance expectations for new technologies and solutions, and working to achieve performance targets.

Requirements

  • B.Sc in Computer Science or Software Engineering or equivalent experience.
  • 15+ years of experience with high-performance networking (RDMA, MPI, NCCL, SHARP).
  • Demonstrated ability in performance evaluation techniques and approaches.
  • Experience with NVIDIA GPUs and the CUDA library. Knowledge of deep learning frameworks like TensorFlow or PyTorch.
  • Expertise in networking collective communication libraries such as NCCL and protocols like RoCE and RDMA.
  • Fast and self-learning capabilities with strong analytical and problem-solving skills.
  • Proficiency in programming languages: Python, Bash, and C++.
  • Experience with a container-based development environment.
  • Great teammate who communicates clearly and works well with others.

Ways to stand out from the crowd

  • Extensive understanding and hands-on experience with AI workloads and benchmarking for distributed LLM training.
  • Knowledge in PyTorch, CUDA, and NCCL libraries.
  • Comprehensive system knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).
  • Strong capabilities in performance evaluation and methods using contemporary tools.

The base salary range is 272,000 USD - 431,250 USD for Level 6, and 320,000 USD - 488,750 USD for Level 7.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 16, 2026.

More jobs at Nvidia

Similar jobs