Senior Accelerated Computing Architect

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 3 Algorithms @ 7 CUDA @ 4 Communication @ 3 Data Structures GPU @ 4 HPC MPI @ 3 Machine Learning OpenCL @ 4 Performance Optimization @ 6 Prioritization @ 6 Profiling @ 4 Python @ 3

Details

NVIDIA is developing software and system architectures for accelerated high-performance computing, scientific computing, machine learning, artificial intelligence, data centers, and automotive computing. This position focuses on advancing accelerated computing through performance optimization, software-hardware co-design, and collaboration across engineering and research teams.

Responsibilities

  • Perform in-depth analysis and optimization to ensure the best possible performance on current and next-generation NVIDIA GPUs.
  • Create and optimize core parallel algorithms, data structures, and reference code for NVIDIA GPUs.
  • Analyze the interplay between hardware and software architectures, core algorithms, programming models, and applications.
  • Collaborate with hardware design, software engineering, product, and research teams to guide the direction of accelerated computing.
  • Investigate accelerated computing applications to facilitate software-hardware co-design.
  • Document and present work through white papers, conference publications, official blog posts, patent applications, and other appropriate materials.

Requirements

  • Master's degree or Ph.D. in Computer Science, Computer Engineering, or Electrical Engineering, or equivalent experience.
  • At least 6 years of relevant work experience.
  • Strong mathematical fundamentals, including linear algebra and numerical methods.
  • Passion for performance optimization.
  • Hands-on experience with massively parallel GPU programming models such as CUDA or OpenCL.
  • Strong knowledge of C and C++, including software design, programming techniques, and algorithms.
  • Experience benchmarking, profiling, and characterizing workloads on GPU and CPU clusters.
  • Good communication and organization skills, with a logical approach to problem solving, time management, and task prioritization.
  • Familiarity with multi-node communication APIs such as MPI, OpenSHMEM, or NVSHMEM is a plus.
  • Familiarity with threading APIs for multicore CPUs and Unix-style inter-process communication APIs is a plus.
  • Familiarity with Python is a plus.

Compensation and Benefits

  • Base salary range of USD 184,000–287,500 for Level 4.
  • Base salary range of USD 224,000–356,500 for Level 5.
  • Eligible for equity and benefits.
  • Salary is determined based on location, experience, and the pay of employees in similar positions.

NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer. NVIDIA uses AI tools in its recruiting processes. Applications will be accepted at least until May 11, 2026.

More jobs at Nvidia

Similar jobs