Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile
Ansible @ 4
CI/CD @ 4
CUDA @ 6
Debugging @ 7
DevOps @ 4
Docker @ 1
GPU @ 4
GitHub @ 1
Java @ 4
JavaScript @ 4
Jenkins @ 4
Kubernetes @ 1
LLM @ 4
Linux @ 7
Mathematics @ 4
NLP @ 4
OpenCL @ 6
Parallel Programming @ 6
PyTorch @ 4
Python @ 4
Slurm @ 1
Software Development @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an experienced software quality assurance engineer to join its platform SWQA team. The role focuses on enterprise server integration, Linux, reliability testing, telemetry, scale-out clusters, test-plan development, AI tools, NLP, DevOps, and CI/CD. The position supports NVIDIA HGX, DGX, and MGX platforms across servers, operating systems, firmware, and the CUDA software stack.
Responsibilities
- Develop and execute platform test plans for NVIDIA HGX, DGX, and MGX platforms, including servers, operating systems, firmware, and the CUDA software stack.
- Install and test operating systems, server firmware, and software stacks.
- Support root-cause analysis of reliability and validation test failures and drive mitigation.
- Build, develop, and debug server- and operating-system-level automation frameworks and tests, including front-end and back-end components.
- Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
- Work in an agile software development team with high production-quality standards.
- Manage the bug lifecycle and collaborate across groups to drive solutions.
Requirements
- Bachelor's degree, or equivalent experience, in a STEM field such as science, technology, engineering, mathematics, or physics.
- Five or more years of proven experience, or a master's degree.
- Experience with operating-system and server-level automation, CI/CD processes, and DevOps using Python, Shell, Ansible, Jenkins, C/C++, Java, and JavaScript.
- Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
- Knowledge of and hands-on experience with model testing, AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, NLP, and LLM benchmarking.
- Experience using AI development tools to create test plans, develop test cases, and automate test cases.
- Experience with firmware, BMC/OpenBMC, network protocols, enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, UEFI specifications, and Redfish is a strong advantage.
- Experience with GitHub, GitLab, Gerrit, PXE, SLURM, Kubernetes, and Docker is a strong advantage.
Preferred Qualifications
- Experience with AI-related tools, LLMs, and NLP.
- Experience working with NVIDIA GPU hardware.
- Understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
- Background in parallel programming, ideally CUDA or OpenCL.
Compensation and Benefits
- Base salary for Level 3: USD 140,000–224,250 per year.
- Base salary for Level 4: USD 168,000–270,250 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until August 23, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
NVIDIA 2027 Internships: Ph.D. Research Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 38-94 per hour
NVIDIA 2027 Internships: Ph.D. Research Computer Vision and Deep Learning
Nvidia · Santa Clara, United States
USD 38-94 per hour
NVIDIA Spring 2027 Internships: Developer and Performance Technology
Nvidia · Santa Clara, United States
USD 20-71 per hour
Deep Learning Compiler Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
NVIDIA 2027 Internships: Ph.D. Research Computer Architecture and Systems
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
NVIDIA 2027 Internships: Software Engineering
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior Deep Learning Frameworks Sustaining Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Software Engineer - Scientific Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year