Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
Algorithms
CUDA @ 4
Communication @ 6
GPU @ 7
HPC
JAX @ 7
LLVM @ 4
Performance Optimization @ 7
Profiling @ 4
PyTorch @ 7
Python @ 7
Rust @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s accelerated computing platform is foundational to modern HPC and AI. CUDA Core Libraries provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software.
The team is hiring a Senior Software Engineer to advance the C++ foundation of CUDA Core Libraries. The role involves designing and optimizing high-performance algorithms and APIs for C++ developers, as well as building foundational libraries, algorithms, and language/runtime infrastructure for CUDA.
Responsibilities
- Design and implement foundational CUDA C++ libraries, parallel algorithms, utilities, and runtime abstractions.
- Compose and optimize GPU algorithms from high-level generic interfaces through low-level implementation.
- Design stable interoperability boundaries that allow core C/C++ functionality to be consumed efficiently from Python and Rust.
- Balance performance, compile time, portability, compatibility, usability, and long-term API evolution.
- Own features throughout their lifecycle, including design, implementation, testing, profiling, benchmarking, documentation, release, and maintenance.
- Improve developer productivity through diagnostics, examples, build integration, tests, benchmarks, and continuous integration.
- Collaborate with Python, Rust, compiler, and runtime engineers during architecture, design, and code reviews.
- Engage with users on performance investigations, API feedback, and correctness issues.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- Eight or more years of relevant software-development experience.
- Strong production programming skills in C and C++, with deep knowledge of modern C++.
- Experience with generic programming, templates, type systems, and standard-library design principles.
- Proven experience developing systems-level software with demanding performance, concurrency, and compatibility requirements.
- Practical experience with CUDA or another parallel or heterogeneous programming environment.
- Experience developing production software or foundational libraries, including testing, profiling, benchmarking, and code review.
- Understanding of API and ABI compatibility and the challenges of exposing C/C++ functionality to other languages.
- Ability to work independently, define project scope, and drive complex work to completion.
- Clear written communication skills for architecture documents, API specifications, and developer documentation.
- Comfort working in large C/C++ codebases with build systems, toolchains, and continuous-integration infrastructure.
Preferred Qualifications
- Strong understanding of CPU/GPU architecture and performance optimization, with hands-on experience in GPU-accelerated stacks such as CUDA C++/Python, PyTorch, JAX, Numba, or CuPy.
- Proficiency with modern C++ and GPU libraries such as Thrust, CUB, and libcudacxx.
- Experience with compiler infrastructure and tooling, including LLVM, Clang, or MLIR.
- Knowledge of binary interfaces, linking, versioning, cross-platform distribution, and interoperability across Python, Rust, and C/C++ stacks.
- Demonstrated interest in developer tools, library design, and improving developer productivity.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The base salary will be determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 13, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.