Senior Compute Kernel Architect, GPU Power

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 7 Communication @ 6 Debugging @ 6 GPU @ 7 HPC

Details

NVIDIA is seeking a Compute Kernel Performance Architect to develop, profile, and analyze CUDA workloads with a strong focus on GPU power behavior. The role involves creating specialized workloads that exercise GPU compute, memory, and I/O subsystems under demanding operating conditions. You will work closely with GPU architects, power architects, silicon validation engineers, and software teams to characterize workload behavior and influence the power architecture of future NVIDIA products.

This position sits at the intersection of GPU architecture, high-performance software, and silicon characterization.

Responsibilities

  • Design and develop CUDA kernels and infrastructure that exercise worst-case power behavior across GPU compute, memory, and I/O subsystems.
  • Profile workloads to understand the relationship between kernel behavior, hardware utilization, performance, and power consumption.
  • Build workloads that generate controlled steady-state and transient power conditions across multiple GPU architectures.
  • Partner with GPU architects and silicon teams to identify functional units and workload patterns that require additional characterization.
  • Support power-stress methodology from pre-silicon modeling and simulation through post-silicon bring-up and validation.

Requirements

  • MS, PhD, or equivalent experience in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
  • 5+ years of experience in CUDA programming, GPU kernel development, high-performance computing, or performance architecture.
  • Hands-on experience developing and optimizing GPU kernels, including work at the PTX or assembly level.
  • Experience with GPU performance-analysis tools such as Nsight Compute, Nsight Systems, nvprof, or equivalent tools.
  • Strong understanding of GPU build principles, including streaming multiprocessors, execution pipelines, memory hierarchy, synchronization, occupancy, and power states.
  • Excellent analytical, debugging, and communication skills.
  • Ability to work effectively across GPU architecture, software, silicon validation, and hardware engineering teams.

Preferred Qualifications

  • Experience crafting GPU power-stress microbenchmarks or test-to-failure workloads.
  • Familiarity with Power Delivery Network concepts, including package- and board-level behavior, impedance, inductance, decoupling, resonance, voltage droop, and overshoot.
  • Understanding of di/dt and how changes in current over time can compose voltage transients.
  • Experience with DVFS, AVFS, clock management, power states, or hardware noise-mitigation mechanisms.
  • Knowledge of how software workload patterns can interact with system-level power-delivery behavior.

Team

The team works at the core of NVIDIA's GPU performance and power stack and collaborates closely with Compute Architecture, Power Architecture, Silicon Solutions, circuit-design teams, and deep-learning software teams. The workloads, tools, and analysis produced by this team help validate current products and influence the design of upcoming NVIDIA GPUs.

Compensation and Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Base salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until July 26, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to fostering an inclusive work environment and equal opportunity employment.

More jobs at Nvidia

Similar jobs