Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 6
CUDA @ 4
Communication @ 6
Debugging
Deep Learning @ 4
Distributed Systems
GPU
LLM @ 4
MPI @ 4
Machine Learning
NCCL @ 4
Networking @ 7
Performance Analysis @ 7
Profiling
PyTorch @ 4
Python @ 6
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing the next era of computing through accelerated computing and artificial intelligence. The AI Networking Codesign and Benchmarking R&D group is seeking a senior software engineer to profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters used for distributed deep learning large language model (LLM) training and inference. The role focuses on collective communication and networking across hardware components including HCAs, switches, CPUs, GPUs, and systems, as well as software layers including LLM applications, machine learning frameworks, communication libraries, and computing libraries.
The engineer will build performance analysis tools and strategies to investigate performance expectations, limitations, and bottlenecks.
Responsibilities
- Characterize AI workloads and deep learning models for large-scale LLM training and inference on NVIDIA supercomputers.
- Work with distributed systems, high-performance networking, and NVIDIA communication libraries.
- Benchmark, profile, and analyze performance to identify bottlenecks, improvements, and optimization opportunities, with a strong emphasis on networking.
- Develop PyTorch trace-based profiling, analysis, and replay tools for benchmarking, debugging, and co-designing network systems for LLM workloads.
- Collaborate with teams across hardware and software to provide performance analysis insights.
- Define performance test plans, set performance expectations for new technologies and solutions, and work toward performance targets.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or equivalent experience.
- 15 or more years of experience with high-performance networking, including RDMA, MPI, NCCL, and SHARP.
- Demonstrated ability with performance evaluation techniques and approaches.
- Experience with NVIDIA GPUs and the CUDA library.
- Knowledge of deep learning frameworks such as TensorFlow or PyTorch.
- Expertise in networking collective communication libraries such as NCCL and protocols such as RoCE and RDMA.
- Strong analytical and problem-solving skills, with the ability to learn quickly and independently.
- Proficiency in Python, Bash, and C++.
- Experience with container-based development environments.
- Clear communication skills and the ability to work effectively as part of a team.
Preferred Qualifications
- Extensive understanding of and hands-on experience with AI workloads and benchmarking for distributed LLM training.
- Knowledge of PyTorch, CUDA, and NCCL libraries.
- Comprehensive system knowledge, including Intel, AMD, and ARM CPUs; NVIDIA GPUs; HCAs; memory; and PCI.
- Strong capabilities in performance evaluation and contemporary performance analysis tools and methods.
Compensation And Benefits
The base salary range is USD 272,000–431,250 for Level 6 and USD 320,000–488,750 for Level 7. Base salary is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until June 16, 2026.