System Software Engineer, Performance - CUDA Driver

at Nvidia
USD 124,000-195,500 per year
MIDDLE
✅ On-site

Tech Stack

AI @ 3 API CUDA @ 3 Communication @ 3 Data Analysis @ 3 Deep Learning @ 3 Experimentation @ 3 GPU @ 5 HPC @ 3 Python @ 3 Robotics @ 3 Software Development @ 3

Details

The AI revolution advances when computation becomes fast, efficient, and economical enough to turn new ideas into products at global scale. CUDA is a critical layer beneath the frameworks, libraries, and applications used across AI, deep learning, high-performance computing, graphics, automotive, robotics, and other CUDA-powered products.

This role focuses on designing and shipping production C/C++ features and optimizations in the CUDA driver and runtime. The work includes tracing workloads across application, operating-system, CPU, interconnect, and GPU boundaries; bringing up new platforms; and using performance evidence to influence future software and hardware direction.

Responsibilities

  • Design, implement, validate, and ship performance-centric features and programming-model capabilities in the CUDA driver and runtime using maintainable, well-tested production C/C++.
  • Optimize critical execution paths, including kernel launch, synchronization, memory management and movement, CPU–GPU coordination, and system interconnect use, for latency, throughput, bandwidth, efficiency, and scalability.
  • Own complex performance problems end to end by understanding workloads, forming hypotheses, creating focused measurements and models, isolating root causes across software and hardware boundaries, implementing production solutions, and validating application-level impact.
  • Establish performance expectations for current and future platforms, characterize new silicon, close software and hardware gaps, and drive performance readiness through product release.
  • Translate workload and platform evidence into CUDA API and programming-model improvements, systems-software direction, and measurement-backed recommendations for future hardware architecture and implementation.
  • Partner with application, library, framework, operating-system, driver, runtime, firmware, GPU architecture, silicon, product, and customer-facing teams.
  • Communicate findings clearly and raise engineering quality through design and code reviews.
  • Lead complex feature development and cross-layer investigations, define performance requirements and technical direction for major subsystems, mentor engineers, and shape hardware/software decisions for future product generations.

Requirements

  • A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
  • At least 2 years of relevant systems-software development experience.
  • Strong production C/C++ systems-programming experience, including delivery of substantial features, optimizations, or production fixes in a complex codebase.
  • Strong operating systems and concurrency foundations, including threads, synchronization, processes, virtual memory, and user/kernel interactions.
  • Strong computer-architecture foundations, including processors, memory hierarchy, caching and coherence, data movement, and system interconnects.
  • Demonstrated success improving real software performance through measurement, identification of limiting mechanisms, implementation of effective solutions, and quantitative validation.
  • Sound technical judgment, ownership of ambiguous problems, and clear communication across organizational and disciplinary boundaries.
  • Direct CUDA or GPU experience is valuable but not required when accompanied by deep systems-software, operating-systems, computer-architecture, and performance-engineering foundations.

Preferred Qualifications

  • Experience developing GPU or accelerator drivers, runtimes, kernel software, firmware, compilers, or other performance-critical low-level systems.
  • Experience with pre-silicon analysis, platform bring-up, performance modeling, or hardware/software co-design.
  • Systems-level performance experience with AI/deep learning, HPC, graphics, automotive, robotics, or similarly demanding workloads.
  • Evidence of technical inventions, such as software-performance patents, novel production designs, or measurement-backed recommendations that influenced a hardware revision or future architecture.
  • Python or another scripting language used for focused experimentation, data analysis, or visualization.

Compensation and Benefits

  • Base salary range: USD 124,000–195,500 per year.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until September 5, 2026.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs