Senior Software QA Test Development Engineer - Diagnostics

at Nvidia
USD 140,000-270,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agile Ansible @ 4 CI/CD @ 4 CUDA @ 6 Debugging @ 7 DevOps @ 4 Docker @ 1 GPU @ 4 GitHub @ 1 Java @ 4 JavaScript @ 4 Jenkins @ 4 Kubernetes @ 1 LLM @ 4 Linux @ 7 Mathematics @ 4 NLP @ 4 OpenCL @ 6 Parallel Programming @ 6 PyTorch @ 4 Python @ 4 Slurm @ 1 Software Development @ 4 TensorFlow @ 4

Details

NVIDIA is seeking an experienced software quality assurance engineer to join its platform SWQA team. The role focuses on enterprise server integration, Linux systems, reliability testing, telemetry, scale-out clusters, test plan development, AI tools, NLP, DevOps, and CI/CD. The position supports NVIDIA HGX, DGX, and MGX platforms across server hardware, operating systems, firmware, and the CUDA software stack.

Responsibilities

  • Develop and execute platform test plans for NVIDIA HGX, DGX, and MGX platforms, including servers, operating systems, firmware, and the CUDA software stack.
  • Install and test operating systems, server firmware, and software stacks.
  • Support root cause analysis for reliability and validation test failures, identifying causes and implementing mitigations.
  • Build, develop, and debug server- and OS-level automation frameworks and tests, including front-end and back-end components.
  • Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
  • Work within an agile software development team with high production-quality standards.
  • Manage the bug lifecycle and collaborate across groups to drive solutions.

Requirements

  • Bachelor's degree or equivalent experience in a STEM field, including science, technology, engineering, mathematics, or physics.
  • At least 5 years of proven experience, or a master's degree.
  • Experience with OS- and server-level automation, CI/CD processes, and DevOps using Python, shell scripting, Ansible, Jenkins, C/C++, Java, and JavaScript.
  • Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
  • Knowledge of and hands-on experience with model testing, AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, as well as NLP and LLM benchmarking.
  • Experience using AI development tools to create test plans, develop test cases, and automate test cases.
  • Experience with firmware, BMC/OpenBMC, network protocols, internal and external enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, UEFI specifications, and Redfish is a significant advantage.
  • Experience with GitHub, GitLab, Gerrit, PXE, SLURM, Stack, Kubernetes, and Docker is a significant advantage.

Preferred Qualifications

  • Experience with AI-related tools, LLMs, and NLP.
  • Experience working with NVIDIA GPU hardware.
  • Solid understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
  • Background in parallel programming, ideally CUDA or OpenCL.

Compensation and Benefits

The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 140,000–224,250 for Level 3 and USD 168,000–270,250 for Level 4. The role is also eligible for equity and benefits.

Applications will be accepted at least until July 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs