System Software Engineer – Data Center Compute Diagnostics

at Nvidia
USD 152,000-241,500 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 4 Communication @ 6 Debugging @ 7 GPU @ 4 Linux @ 6 NCCL NVLink Networking @ 4 PyTorch Python @ 7

Details

NVIDIA is seeking a system software engineer to develop low-level diagnostic software for next-generation data center GPUs and rack-scale AI systems. The team builds software that exercises and validates complex hardware, including compute engines, memory and cache subsystems, NICs, PCIe and NVLink interfaces, power delivery, and thermal behavior.

This role is well suited to an embedded, firmware, device-driver, hardware-validation, or systems software engineer who enjoys working close to hardware. Relevant experience may come from GPUs, CPUs, networking, storage, servers, embedded systems, or other complex silicon-based products. Prior GPU, CUDA, or GEMM experience is helpful but not required. The engineer will own well-scoped components of the diagnostic software from design through implementation, validation, productization, and field support, while collaborating with hardware architects, driver developers, silicon-validation engineers, manufacturing teams, and field engineers.

Responsibilities

  • Develop diagnostic and stress software in C, C++, and Python for complex hardware systems.
  • Collaborate with hardware blocks, firmware, Linux device drivers, registers, telemetry, and low-level debugging tools.
  • Bring up and validate new silicon and system features using pre-production hardware and software.
  • Create targeted tests for compute engines, memory and cache subsystems, DMA engines, PCIe and NVLink interfaces, power, and thermal behavior.
  • Investigate hardware and software failures involving memory errors, ECC, data integrity, performance, thermals, voltage/frequency behavior, and high-speed interfaces.
  • Contribute to diagnostic and stress workloads ranging from low-level GPU hardware tests to higher-level AI workloads.
  • Develop expertise in CUDA programming, GEMM-style compute, NCCL, and PyTorch-based workloads.
  • Use modern development and analysis tools, including AI-assisted tools where appropriate, to accelerate coding, debugging, test creation, and failure analysis.

Requirements

  • BS or MS degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent experience.
  • 5+ years of experience in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, silicon bring-up, or related fields.
  • Strong programming skills in C and C++, plus working proficiency in Python.
  • Experience developing software that interacts with hardware, firmware, device drivers, hardware registers, or low-level interfaces.
  • Experience creating diagnostics, validation tests, stress tests, manufacturing tests, or other software used to isolate hardware or system failures.
  • Understanding of computer architecture concepts such as memory, caches, interrupts, DMA, buses, and device I/O.
  • Strong debugging and problem-solving skills, including the ability to investigate failures across hardware and software boundaries.
  • Ability to take ownership of a well-scoped problem and drive it to completion while collaborating with a technical lead and multifunctional teams.
  • Good written and verbal communication skills.

Benefits

  • Base salary range of USD 152,000–241,500 per year, determined by location, experience, and compensation for similar positions.
  • Eligibility for equity and benefits.
  • NVIDIA offers a comprehensive benefits package.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until August 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs