Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 3
Communication @ 4
Docker @ 3
GPU
HPC @ 4
Kubernetes @ 3
MPI @ 4
NCCL @ 4
Networking
Parallel Programming @ 4
Python @ 6
Slurm @ 3
System Architecture @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
- Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack.
- Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available.
- Triage and root-cause performance issues reported by our customers.
- Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information.
- Collaborate with a very dynamic team across multiple time zones.
Requirements
- M.S. (or equivalent experience) or PhD in Computer Science, or related field with relevant performance engineering and HPC experience.
- 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM).
- Experience conducting performance benchmarking and triage on large scale HPC clusters.
- Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals).
- Implement micro-benchmarks in C/C++, read and modify the code base when required.
- Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python.
- Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker).
- Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones.
Benefits
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Senior Software Engineer, Compute Sanitizer - GPU
Nvidia · United States
USD 184,000-356,500 per year
Senior Backend Platform Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Applied AI Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Systems Software Engineer, Accelerated Kubernetes Performance And Scale - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Senior Hpc Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Architect - Deep Learning And Hpc Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Ai And Ml Infra Software Engineer, GPU Clusters
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
AI/Ml Specialist Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year