Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Bash @ 4
Communication @ 7
Deep Learning @ 4
GPU
HPC @ 7
Linux
Machine Learning
Mathematics @ 4
Python @ 4
System Administration
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters.
Be a key player to enable the most exciting computing hardware and software and contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on and implement at-scale system administration and tuning mechanisms for large-scale compute runs. You will work with the latest accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.
What you’ll be doing
- Provide engineering solutions to operationalize the latest GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assisting systems, machine learning/deep learning engineers in building creative solutions based on NVIDIA technology.
- Be an internal reference for system administration, at-scale system analysis, and other datacenter and large-scale GPU-accelerated system solutions among the NVIDIA technical community.
What we need to see
- 8+ years of experience using in accelerated computing for datacenter/HPC-based Enterprise computing solutions.
- Solid understanding of accelerated computing scheduling and I/O stacks.
- C/C++/Python/Bash programming/scripting experience.
- Experience working with engineering or academic research community supporting high performance computing or deep learning.
- Experience with parallel filesystems.
- Strong teamwork and communication skills, both verbal and written.
- Ability to multitask effectively in a dynamic environment.
- Action driven with strong analytical skills.
- Desire to be involved in multiple diverse and innovative projects.
- BS (or equivalent experience) in Engineering, Mathematics, Physics, or Computer Science. MS or PhD desirable.
Ways to stand out from the crowd
- Deep Learning framework skills.
- Exposure to using and deploying telemetry and visualization pipelines
- Exposure to container technology and Linux performance tools.
Additional information
- Base salary range: 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
- You will also be eligible for equity and benefits.
- Applications accepted at least until July 22, 2026.