Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
CUDA
GPU @ 7
HPC @ 4
Kubernetes @ 3
Linux @ 4
Python @ 7
Reporting @ 4
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior System Architect to solve a complex challenge in accelerated computing: failure attribution at scale. As EDA or equivalent workloads scale across thousands of heterogeneous nodes, a single failure can cause significant resource waste. The engineer will develop and build an automated framework that ingests telemetry from CPU and GPU clusters to identify the root cause of job failures in real time, distinguishing between hardware faults, infrastructure instability, and software defects.
Responsibilities
- Build scalable failure attribution frameworks, including a high-fidelity "flight recorder" for EDA jobs that captures CPU, GPU, and fabric state at the moment of failure.
- Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions with system-level events such as OOM kills and NUMA-related hangs.
- Implement low-overhead distributed logging and tracing mechanisms using tracing tools or custom agents across multi-node Slurm or Kubernetes clusters.
- Develop heuristics and machine-learning models to classify failures as hardware faults, software bugs, or environment issues, reducing Mean Time to Identify (MTTI) for R&D teams.
- Work with hardware and infrastructure teams to define signals of impending failure and enable proactive job migration or checkpointing before crashes occur.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science or Electrical Engineering, or equivalent experience, with 6 or more years of systems programming experience.
- Experience building automated Root Cause Analysis (RCA) pipelines for HPC or cloud-scale environments.
- Expert knowledge of x86/ARM node-level metrics, including IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
- Strong C++ and Python programming skills, including the ability to build high-performance daemons that monitor system health without affecting workload performance.
- Familiarity with cluster resource managers such as Slurm, LSF, or Kubernetes, including job lifecycle and signal propagation management.
Preferred Qualifications
- Expert knowledge of the Linux kernel and error-reporting interfaces such as
/dev/mcelog,dmesg, andjournald. - Understanding of how the Linux kernel handles hardware exceptions and memory faults.
- Deep experience with NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state dumps.
- Experience with non-intrusive application-health monitoring and syscall-level failure-pattern analysis.
- Experience with checkpoint/restore technologies such as CRIU and their application in long-running EDA flows.
Compensation and Benefits
- Base salary range: $184,000–$287,500 for Level 4.
- Base salary range: $224,000–$356,500 for Level 5.
- Base salary is determined by location, experience, and compensation for similar positions.
- Eligible for equity and benefits.
- This is a hybrid position.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior HPC AI Cluster Engineer
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
NVIDIA Spring 2027 Internships: Developer and Performance Technology
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year