Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Ansible @ 3
CUDA @ 3
Communication @ 6
Debugging @ 4
Deep Learning @ 4
Docker @ 3
GPU
HPC @ 4
InfiniBand @ 4
Kubernetes @ 3
MPI @ 6
NCCL @ 6
NVLink
Networking
Parallel Programming @ 6
PyTorch @ 4
Python @ 6
Slurm @ 3
System Architecture @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's GPU Communications Libraries and Networking team develops libraries such as NCCL, NVSHMEM, and UCX for deep learning and high-performance computing. The team is seeking a performance engineer to help shape the roadmap of communication libraries used by applications running across large-scale, multi-GPU and multi-node environments.
GPUs are connected through high-speed interconnects such as NVLink and PCIe within a node, and through InfiniBand and Ethernet across nodes. Communication performance directly affects end-to-end application performance, particularly at scales involving tens of thousands of GPUs.
Responsibilities
- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
- Study the interaction of communication libraries with GPU, CPU, networking, and other hardware and software components.
- Evaluate proof-of-concepts and conduct trade-off analyses when multiple solutions are available.
- Triage and identify the root causes of performance issues reported by customers.
- Collect performance data and build tools and infrastructure to visualize and analyze it.
- Collaborate with a dynamic team across multiple time zones.
Requirements
- M.S. degree or equivalent experience, or PhD, in Computer Science or a related field, with relevant performance engineering and HPC experience.
- At least 3 years of experience with parallel programming and at least one communication runtime, such as MPI, NCCL, UCX, or NVSHMEM.
- Experience conducting performance benchmarking and triage on large-scale HPC clusters.
- Good understanding of computer system architecture, hardware-software interactions, operating system principles, and systems software fundamentals.
- Ability to implement micro-benchmarks in C or C++ and read and modify code when required.
- Ability to debug performance issues across the hardware and software stack.
- Proficiency in a scripting language, preferably Python.
- Familiarity with containers, cloud provisioning, and scheduling tools such as Kubernetes, SLURM, Ansible, and Docker.
- Adaptability, willingness to learn new areas and tools, and the ability to communicate effectively across teams and time zones.
Preferred Qualifications
- Practical experience with InfiniBand or Ethernet networks, including RDMA, topologies, and congestion control.
- Experience debugging network issues in large-scale deployments.
- Familiarity with CUDA programming and/or GPUs.
- Experience with deep learning frameworks such as PyTorch and TensorFlow.
Compensation and Benefits
- Base salary range of $152,000–$241,500 for Level 3.
- Base salary range of $184,000–$287,500 for Level 4.
- Base salary is determined by location, experience, and compensation for employees in similar positions.
- Eligible employees may also receive equity and benefits.
Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive, equal-opportunity work environment.