Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile
Ansible @ 4
CI/CD @ 4
CUDA @ 6
Debugging @ 7
DevOps @ 4
Docker @ 1
GPU @ 4
GitHub @ 1
HPC
Java @ 4
JavaScript @ 4
Jenkins @ 4
Kubernetes @ 1
LLM @ 4
Linux @ 4
Mathematics @ 4
NLP @ 4
Networking
OpenCL @ 6
Parallel Programming @ 6
PyTorch @ 4
Python @ 4
Slurm @ 1
Software Development @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is the world leader in GPU computing, serving markets including gaming, automotive, vision, high-performance computing, data centers, networking, and AI computing. The Platform Software Quality Assurance team is seeking an experienced engineer with enterprise server integration, Linux, reliability testing, scale-out clusters, test plan development, AI tools, NLP, DevOps, and CI/CD experience.
Responsibilities
- Develop and execute NVIDIA HGX, DGX, and MGX platform test plans covering servers, operating systems, firmware, and the CUDA software stack from design documents.
- Install and test various operating systems, server firmware, and software stacks.
- Support root cause analysis for reliability and validation test failures, identifying causes and achieving mitigation.
- Build, develop, and debug server- and operating-system-level automation frameworks and tests, including front-end and back-end components.
- Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
- Work in an agile software development team with high production-quality standards.
- Manage the bug lifecycle and collaborate across groups to drive solutions.
Requirements
- Bachelor's degree, or equivalent experience, in a STEM field such as science, technology, engineering, mathematics, or physics.
- Five or more years of proven experience, or a master's degree.
- Experience with operating-system and server-level automation, CI/CD processes, and DevOps using Python, Shell, Ansible, Jenkins, C/C++, Java, and JavaScript.
- Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
- Knowledge of and hands-on experience with model testing, AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, as well as NLP and LLM benchmarking.
- Experience using AI development tools to create test plans, develop test cases, and automate test cases.
- Experience with firmware, BMC/OpenBMC, network protocols, enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, and UEFI specifications is a strong plus.
- Experience with GitHub, GitLab, Gerrit, PXE, SLURM, Kubernetes, and Docker is a strong plus.
Preferred Qualifications
- Experience with AI-related tools, LLMs, and NLP.
- Experience working with NVIDIA GPU hardware.
- Understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
- Background in parallel programming, ideally CUDA or OpenCL.
Compensation and Benefits
- Base salary range for Level 3: USD 140,000–224,250 per year.
- Base salary range for Level 4: USD 168,000–270,250 per year.
- Compensation is determined based on location, experience, and the pay of employees in similar positions.
- Eligible employees may also receive equity and benefits.
- Applications will be accepted at least until August 15, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior System Software Engineer, Platform - OpenBMC
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Security Engineer, RTOS and Virtualization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVIDIA Warp
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software And System Architect
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Compiler Engineer Infrastructure
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Frameworks Sustaining Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year