Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
CI/CD @ 4
CUDA @ 4
Communication @ 7
Computer Vision
Debugging @ 7
Deep Learning @ 4
GPU @ 4
HPC
Jira @ 4
LLM
MPI @ 4
Mathematics @ 4
Parallel Programming @ 7
Performance Optimization @ 4
Product Management @ 4
Project Management @ 4
PyTorch @ 4
Python @ 6
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for software engineers to contribute to the design and development of libraries and tools that simplify and accelerate computing for unstructured sparsity in deep learning and high-performance computing. The team develops GPU-accelerated libraries and SDKs used for applications including large language models, computer-aided engineering, quantum chemistry, autonomous vehicles, computer vision, data analytics, and scientific and engineering simulations.
The role involves developing solutions for generalized sparse tensor computations, domain-specific language specifications for sparse storage formats, and on-demand code generation. Candidates should have experience developing accelerated computing software and be motivated to advance accelerated computing across deep learning frameworks such as PyTorch.
Responsibilities
- Design and develop a C++-based system to simplify and accelerate computing for unstructured sparsity in deep learning and high-performance computing on NVIDIA GPUs.
- Enable the system in commonly used deep learning languages and frameworks, including Python and PyTorch.
- Evaluate and improve system performance on real-life applications.
- Improve library quality, performance, and maintainability by writing effective, well-tested production code.
- Collaborate with product management and internal and external partners to understand feature and performance requirements and contribute to technical roadmaps.
Requirements
- Bachelor's, master's, or PhD degree in Computer Science, Applied Mathematics, or a related field, or equivalent experience.
- At least 6 years of experience developing, debugging, and optimizing high-performance software using C++ and parallel programming. Experience with sparse linear algebra applications and technologies such as CUDA, MPI, or OpenMP is preferred.
- Experience with domain-specific language design and compiler optimizations, particularly sparse compilers such as MLIR or TACO.
- Excellent C++, Python, and CUDA programming skills.
- Strong collaboration, communication, and documentation skills. Experience working in a globally distributed organization is preferred.
Preferred Qualifications
- Strong understanding of sparse computations, particularly sparsity in AI and high-performance computing.
- Good understanding of large language models, deep learning methods, and deep learning frameworks.
- Experience with low-level GPU performance optimization.
- Understanding of numerical linear algebra methods, including direct and iterative solvers.
- Experience adopting and advancing software development practices such as CI/CD, and using project management tools such as JIRA.
Compensation and Benefits
- Base salary range: $184,000-$287,500 USD for Level 4.
- Base salary range: $224,000-$356,500 USD for Level 5.
- Salary is determined based on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
Applications will be accepted at least until June 16, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.