Distinguished Resiliency and Safety Architect, GPU Diagnostics

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 6 Compliance Debugging @ 7 Deep Learning @ 3 GPU @ 3 HPC Machine Learning @ 3 Networking Python @ 4 Robotics Software Development @ 4 System Architecture @ 4

Details

NVIDIA is seeking a Resiliency and Safety Architect to support the development of GPU diagnostics for resiliency in data centers and functional safety in autonomous vehicles and robots. The role focuses on NVIDIA GPUs and systems-on-chip powering artificial intelligence, self-driving, robotics, and industrial applications.

Responsibilities

  • Design, develop, and maintain a diagnostics software suite to stress-test NVIDIA GPUs and SoCs and identify hardware defects, including defects that cause silent data corruption.
  • Develop tests for large-scale data center GPU and safety SoC deployments in package, board, and rack configurations spanning GPUs, CPUs, and networking SoCs.
  • Address coverage gaps identified through silicon failures on customer workloads or test suites.
  • Enhance diagnostics to improve failure repeatability and optimize test time.
  • Develop low-level GPU tests for automotive functional safety, including routines that exercise instruction sets, memory subsystems, and interrupt mechanisms in compliance with ISO 26262 and related safety standards.
  • Collaborate with architecture, RTL, and verification teams to ensure safety coverage, correctness, and robustness across GPU generations.
  • Investigate silent data corruption, intermittent faults, and difficult-to-reproduce field failures, including customer returns, to determine root causes and improve diagnostic detection.
  • Support deployment of diagnostics in pre-production qualification environments and large-scale production use.

Requirements

  • Master's or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field, or equivalent experience.
  • At least 15 years of relevant experience.
  • Ability to reason across hardware and software boundaries to debug complex system-level issues.
  • In-depth understanding of high-performance computing system architecture and microarchitecture.
  • Strong knowledge of hardware failure mechanisms that can result in incorrect computation.
  • Proficiency in C/C++ and CUDA programming.
  • Experience with Python or similar scripting and automation tools.
  • Understanding of the software development life cycle, from requirements through testing closure and maintenance, including customer releases and documentation.
  • Excellent interpersonal and collaboration skills for working with on-site and remote teams.
  • Strong debugging and analytical skills.
  • Self-driven and results-oriented.

Preferred Qualifications

  • Familiarity with GPU and SoC architectures and machine learning or deep learning concepts.
  • Understanding of factors that cause silent data corruption in hardware.
  • Ability to use high-performance libraries and write hand-crafted kernels to create stress conditions that induce hardware failures.
  • Experience in embedded software development.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

Applications will be accepted at least until February 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs