Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Agile
Ansible @ 4
CI/CD @ 7
CUDA @ 6
Debugging @ 7
Deep Learning
DevOps @ 7
Docker @ 1
GPU @ 4
GitHub @ 1
HPC
Java @ 4
JavaScript @ 4
Jenkins @ 4
Kubernetes @ 1
LLM @ 4
Linux @ 7
Mathematics @ 4
NLP @ 7
Networking
OpenCL @ 6
Parallel Programming @ 6
PyTorch @ 4
Python @ 4
Slurm @ 1
Software Development @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an outstanding individual to join its platform software quality assurance team. The role requires enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale-out cluster testing, test plan development, AI tools and NLP experience, and DevOps and CI/CD expertise. NVIDIA develops GPU computing platforms supporting gaming, automotive, vision, high-performance computing, data centers, networking, deep learning, analytics, and autonomous vehicles.
Responsibilities
- Develop and execute NVIDIA HGX, DGX, and MGX platform test plans covering servers, operating systems, firmware, and the CUDA software stack based on design documentation.
- Install and test operating systems, server firmware, and software stacks.
- Support root-cause analysis for reliability and validation test failures, identify root causes, and achieve mitigation.
- Build, develop, and debug server- and operating-system-level automation frameworks and tests, including front-end and back-end components.
- Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
- Work in an agile software development team with high production-quality standards.
- Manage the bug lifecycle and collaborate across groups to drive solutions.
Requirements
- Bachelor's degree, or equivalent experience, in a STEM field such as science, technology, engineering, mathematics, or physics.
- Five or more years of proven experience, or a master's degree.
- Experience with operating-system and server-level automation, CI/CD processes, and DevOps using Python, shell scripting, Ansible, Jenkins, C/C++, Java, and JavaScript.
- Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
- Hands-on experience with model testing, AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, as well as NLP and LLM benchmarking.
- Experience using AI development tools to create test plans, develop test cases, and automate test cases.
- Experience with firmware, BMC/OpenBMC, network protocols, internal and external enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, and UEFI specifications is a strong advantage.
- Experience with GitHub, GitLab, Gerrit, PXE, SLURM, Kubernetes, and Docker is a strong advantage.
Preferred Qualifications
- Experience with AI-related tools, LLMs, and NLP.
- Experience working with NVIDIA GPU hardware.
- Solid understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
- Background in parallel programming, ideally CUDA or OpenCL.
Benefits
The role offers competitive salaries, equity, and a generous benefits package. The base salary is determined by location, experience, and compensation for similar positions. Applications will be accepted at least until July 1, 2026. This posting is for an existing vacancy.