Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Ansible @ 3
CUDA @ 3
Communication @ 6
Debugging @ 4
Deep Learning @ 4
Docker @ 3
GPU
HPC @ 4
InfiniBand @ 4
Kubernetes @ 3
MPI @ 6
NCCL @ 6
NVLink
Networking
Parallel Programming @ 6
PyTorch @ 4
Python @ 6
Slurm @ 3
System Architecture @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is leading the way in groundbreaking developments in artificial intelligence, high-performance computing, and visualization. The GPU, NVIDIA's invention, serves as the visual cortex of modern computers and is at the heart of its products and services.
The GPU communication libraries NCCL, NVSHMEM, and GPUDirect are crucial for scaling deep learning and HPC applications. These applications run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects such as NVLink and PCIe within a node, and with high-speed networking such as InfiniBand and Ethernet across nodes. Communication performance between GPUs has a direct impact on end-to-end application performance, particularly at large scale.
Responsibilities
- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
- Study the interaction of communication libraries with hardware and software components across the stack, including GPUs, CPUs, and networking.
- Evaluate proofs of concept and conduct trade-off analyses when multiple solutions are available.
- Triage and identify the root causes of performance issues reported by customers.
- Collect performance data and build tools and infrastructure to visualize and analyze the information.
- Collaborate with a dynamic team across multiple time zones.
Requirements
- M.S. or Ph.D. in Computer Science or a related field, or equivalent experience, with relevant performance engineering and HPC experience.
- At least 3 years of experience with parallel programming and at least one communication runtime, such as MPI, NCCL, UCX, or NVSHMEM.
- Experience conducting performance benchmarking and triage on large-scale HPC clusters.
- Good understanding of computer system architecture, hardware-software interactions, operating-system principles, and systems software fundamentals.
- Ability to implement micro-benchmarks in C or C++ and read and modify code when required.
- Ability to debug performance issues across the entire hardware and software stack.
- Proficiency in a scripting language, preferably Python.
- Familiarity with containers, cloud provisioning, and scheduling tools such as Kubernetes, SLURM, Ansible, and Docker.
- Adaptability, a passion for learning new areas and tools, and the ability to work and communicate effectively across teams and time zones.
Preferred Qualifications
- Practical experience with InfiniBand or Ethernet networks, including RDMA, network topologies, and congestion control.
- Experience debugging network issues in large-scale deployments.
- Familiarity with CUDA programming and/or GPUs.
- Experience with deep learning frameworks such as PyTorch and TensorFlow.
Benefits
NVIDIA offers highly competitive salaries, an extensive benefits package, and a work environment that promotes diversity, inclusion, and flexibility. NVIDIA is an equal opportunity employer committed to fostering a supportive and empowering workplace.
Compensation
For Poland, the base salary range is 221,250 PLN–383,500 PLN for Level 3 and 292,500 PLN–507,000 PLN for Level 4. The base salary is determined based on location, experience, and the pay of employees in similar positions.