Distinguished Resiliency and Safety Architect, GPU Diagnostics
at Nvidia
USD 320,000-488,800 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 6
Compliance
Debugging @ 7
Deep Learning @ 3
GPU @ 3
HPC
Machine Learning @ 3
Networking
Python @ 4
Robotics
Software Development @ 4
System Architecture @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Resiliency and Safety Architect to support the development of GPU diagnostics for resiliency in data centers and functional safety in autonomous vehicles and robots. The role focuses on NVIDIA GPUs and systems-on-chip powering artificial intelligence, self-driving, robotics, and industrial applications.
Responsibilities
- Design, develop, and maintain a diagnostics software suite to stress-test NVIDIA GPUs and SoCs and identify hardware defects, including defects that cause silent data corruption.
- Develop tests for large-scale data center GPU and safety SoC deployments in package, board, and rack configurations spanning GPUs, CPUs, and networking SoCs.
- Address coverage gaps identified through silicon failures on customer workloads or test suites.
- Enhance diagnostics to improve failure repeatability and optimize test time.
- Develop low-level GPU tests for automotive functional safety, including routines that exercise instruction sets, memory subsystems, and interrupt mechanisms in compliance with ISO 26262 and related safety standards.
- Collaborate with architecture, RTL, and verification teams to ensure safety coverage, correctness, and robustness across GPU generations.
- Investigate silent data corruption, intermittent faults, and difficult-to-reproduce field failures, including customer returns, to determine root causes and improve diagnostic detection.
- Support deployment of diagnostics in pre-production qualification environments and large-scale production use.
Requirements
- Master's or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field, or equivalent experience.
- At least 15 years of relevant experience.
- Ability to reason across hardware and software boundaries to debug complex system-level issues.
- In-depth understanding of high-performance computing system architecture and microarchitecture.
- Strong knowledge of hardware failure mechanisms that can result in incorrect computation.
- Proficiency in C/C++ and CUDA programming.
- Experience with Python or similar scripting and automation tools.
- Understanding of the software development life cycle, from requirements through testing closure and maintenance, including customer releases and documentation.
- Excellent interpersonal and collaboration skills for working with on-site and remote teams.
- Strong debugging and analytical skills.
- Self-driven and results-oriented.
Preferred Qualifications
- Familiarity with GPU and SoC architectures and machine learning or deep learning concepts.
- Understanding of factors that cause silent data corruption in hardware.
- Ability to use high-performance libraries and write hand-crafted kernels to create stress conditions that induce hardware failures.
- Experience in embedded software development.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
Applications will be accepted at least until February 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Performance Compiler Engineer - Triton
Nvidia · Redmond, United States
USD 184,000-287,500 per year
Senior Systems Performance Engineer
Nvidia · Santa Clara, United States
USD 136,000-258,800 per year
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software SDET Test Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year