Principal Developer, AI Networking

at Nvidia
USD 272,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Bash @ 6 CUDA @ 4 Communication @ 6 Debugging Deep Learning @ 4 Distributed Systems GPU LLM @ 4 MPI @ 4 Machine Learning NCCL @ 4 Networking @ 7 Performance Analysis @ 7 Profiling PyTorch @ 4 Python @ 6 TensorFlow @ 4

Details

NVIDIA is developing the next era of computing through accelerated computing and artificial intelligence. The AI Networking Codesign and Benchmarking R&D group is seeking a senior software engineer to profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters used for distributed deep learning large language model (LLM) training and inference. The role focuses on collective communication and networking across hardware components including HCAs, switches, CPUs, GPUs, and systems, as well as software layers including LLM applications, machine learning frameworks, communication libraries, and computing libraries.

The engineer will build performance analysis tools and strategies to investigate performance expectations, limitations, and bottlenecks.

Responsibilities

  • Characterize AI workloads and deep learning models for large-scale LLM training and inference on NVIDIA supercomputers.
  • Work with distributed systems, high-performance networking, and NVIDIA communication libraries.
  • Benchmark, profile, and analyze performance to identify bottlenecks, improvements, and optimization opportunities, with a strong emphasis on networking.
  • Develop PyTorch trace-based profiling, analysis, and replay tools for benchmarking, debugging, and co-designing network systems for LLM workloads.
  • Collaborate with teams across hardware and software to provide performance analysis insights.
  • Define performance test plans, set performance expectations for new technologies and solutions, and work toward performance targets.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, or equivalent experience.
  • 15 or more years of experience with high-performance networking, including RDMA, MPI, NCCL, and SHARP.
  • Demonstrated ability with performance evaluation techniques and approaches.
  • Experience with NVIDIA GPUs and the CUDA library.
  • Knowledge of deep learning frameworks such as TensorFlow or PyTorch.
  • Expertise in networking collective communication libraries such as NCCL and protocols such as RoCE and RDMA.
  • Strong analytical and problem-solving skills, with the ability to learn quickly and independently.
  • Proficiency in Python, Bash, and C++.
  • Experience with container-based development environments.
  • Clear communication skills and the ability to work effectively as part of a team.

Preferred Qualifications

  • Extensive understanding of and hands-on experience with AI workloads and benchmarking for distributed LLM training.
  • Knowledge of PyTorch, CUDA, and NCCL libraries.
  • Comprehensive system knowledge, including Intel, AMD, and ARM CPUs; NVIDIA GPUs; HCAs; memory; and PCI.
  • Strong capabilities in performance evaluation and contemporary performance analysis tools and methods.

Compensation And Benefits

The base salary range is USD 272,000–431,250 for Level 6 and USD 320,000–488,750 for Level 7. Base salary is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until June 16, 2026.

More jobs at Nvidia

Similar jobs