Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
Communication @ 6
Debugging @ 7
GPU @ 4
Linux @ 6
NCCL
NVLink
Networking @ 4
PyTorch
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a system software engineer to develop low-level diagnostic software for next-generation data center GPUs and rack-scale AI systems. The team builds software that exercises and validates complex hardware, including compute engines, memory and cache subsystems, NICs, PCIe and NVLink interfaces, power delivery, and thermal behavior.
This role is well suited to an embedded, firmware, device-driver, hardware-validation, or systems software engineer who enjoys working close to hardware. Relevant experience may come from GPUs, CPUs, networking, storage, servers, embedded systems, or other complex silicon-based products. Prior GPU, CUDA, or GEMM experience is helpful but not required. The engineer will own well-scoped components of the diagnostic software from design through implementation, validation, productization, and field support, while collaborating with hardware architects, driver developers, silicon-validation engineers, manufacturing teams, and field engineers.
Responsibilities
- Develop diagnostic and stress software in C, C++, and Python for complex hardware systems.
- Collaborate with hardware blocks, firmware, Linux device drivers, registers, telemetry, and low-level debugging tools.
- Bring up and validate new silicon and system features using pre-production hardware and software.
- Create targeted tests for compute engines, memory and cache subsystems, DMA engines, PCIe and NVLink interfaces, power, and thermal behavior.
- Investigate hardware and software failures involving memory errors, ECC, data integrity, performance, thermals, voltage/frequency behavior, and high-speed interfaces.
- Contribute to diagnostic and stress workloads ranging from low-level GPU hardware tests to higher-level AI workloads.
- Develop expertise in CUDA programming, GEMM-style compute, NCCL, and PyTorch-based workloads.
- Use modern development and analysis tools, including AI-assisted tools where appropriate, to accelerate coding, debugging, test creation, and failure analysis.
Requirements
- BS or MS degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent experience.
- 5+ years of experience in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, silicon bring-up, or related fields.
- Strong programming skills in C and C++, plus working proficiency in Python.
- Experience developing software that interacts with hardware, firmware, device drivers, hardware registers, or low-level interfaces.
- Experience creating diagnostics, validation tests, stress tests, manufacturing tests, or other software used to isolate hardware or system failures.
- Understanding of computer architecture concepts such as memory, caches, interrupts, DMA, buses, and device I/O.
- Strong debugging and problem-solving skills, including the ability to investigate failures across hardware and software boundaries.
- Ability to take ownership of a well-scoped problem and drive it to completion while collaborating with a technical lead and multifunctional teams.
- Good written and verbal communication skills.
Benefits
- Base salary range of USD 152,000–241,500 per year, determined by location, experience, and compensation for similar positions.
- Eligibility for equity and benefits.
- NVIDIA offers a comprehensive benefits package.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until August 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.