Senior Software Engineer, CUDA Python Core Libraries

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 4 Algorithms @ 4 CUDA @ 4 Communication @ 6 GPU @ 4 HPC JAX @ 7 LLVM @ 4 Performance Optimization @ 7 Profiling @ 4 PyTorch @ 7 Python @ 4 Rust @ 6

Details

NVIDIA’s accelerated computing platform is foundational to modern HPC and AI. The CUDA Core Libraries enable developers to build fast, reliable, and scalable GPU-accelerated software.

This role focuses on advancing the Python experience for CUDA Core Libraries by building Pythonic APIs, language bindings, algorithms, and runtime infrastructure on top of native C/C++ foundations. The position involves developing foundational libraries, algorithms, and language/runtime infrastructure for developers and AI coding agents.

Responsibilities

  • Design and implement idiomatic Python APIs and bindings for foundational CUDA capabilities and GPU algorithms.
  • Develop and integrate native C/C++ components supporting Python-facing functionality.
  • Define reliable and efficient interoperability boundaries between Python, C/C++, Rust, and other languages.
  • Develop high-performance interfaces that minimize Python and native-language integration overhead.
  • Own features throughout their lifecycle, including design, implementation, testing, profiling, benchmarking, documentation, release, and long-term maintenance.
  • Improve the Python developer experience through typing, packaging, examples, diagnostics, continuous integration, and compatibility testing.
  • Collaborate with C/C++, Rust, compiler, and runtime engineers on shared architecture and API decisions.
  • Work directly with users to investigate correctness, usability, compatibility, and performance issues.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • At least 8 years of relevant software-development experience.
  • Strong production programming skills in both Python and C/C++.
  • Experience building Python interfaces to native or systems-level software.
  • Understanding of systems software concepts, performance, concurrency, and API design.
  • Practical experience with parallel, heterogeneous, or GPU programming.
  • Experience developing production software or widely used libraries, including testing, profiling, benchmarking, packaging, and code review.
  • Ability to work independently, define project scope, and drive complex work to completion.
  • Clear written communication skills for API specifications, technical designs, and user documentation.
  • Comfort working in large codebases spanning Python, C/C++, build systems, packaging, and continuous-integration infrastructure.

Preferred Qualifications

  • Strong understanding of CPU/GPU architecture and performance optimization, with hands-on experience in GPU-accelerated stacks such as CUDA C++/Python, PyTorch, JAX, Numba, or CuPy.
  • Proficiency with modern C++ and GPU libraries such as Thrust, CUB, and libcudacxx.
  • Experience with compiler infrastructure and tooling, including LLVM, Clang, or MLIR.
  • Expertise in designing low-overhead interoperability between Python and native languages, including exposure to Rust in mixed-language stacks.
  • Demonstrated interest in developer tools, library design, and improving developer productivity.

Compensation and Benefits

The base salary is determined based on location, experience, and the pay of employees in similar positions. The base salary ranges are 184,000 USD–287,500 USD for Level 4 and 224,000 USD–356,500 USD for Level 5. Employees are also eligible for equity and benefits.

Applications will be accepted at least until July 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs