Senior Software Development Engineer in Test - Datacenter Server OS
at Nvidia
USD 140,000-270,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile
Ansible @ 4
CI/CD @ 4
CUDA @ 6
Debugging @ 7
DevOps @ 4
Docker @ 1
GPU @ 4
GitHub @ 1
Java @ 4
JavaScript @ 4
Jenkins @ 4
Kubernetes @ 1
LLM @ 4
Linux @ 7
Mathematics @ 4
NLP @ 4
OpenCL @ 6
Parallel Programming @ 6
PyTorch @ 4
Python @ 4
Slurm @ 1
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an outstanding software quality assurance professional to join its platform SWQA team. The role focuses on enterprise server integration, Linux, reliability testing, scale-out clusters, test plan development, AI tools and NLP, DevOps, and CI/CD for NVIDIA HGX, DGX, and MGX platforms.
Responsibilities
- Develop and execute test plans for NVIDIA HGX, DGX, and MGX platforms covering servers, operating systems, firmware, and the CUDA software stack based on design documentation.
- Install and test operating systems, server firmware, and software stacks.
- Support root cause analysis for reliability and validation test failures and drive mitigation.
- Build, develop, and debug server- and operating-system-level automation frameworks and tests.
- Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
- Work in an agile software development team with high production-quality standards.
- Manage the bug lifecycle and collaborate across groups to drive solutions.
Requirements
- Bachelor's degree, or equivalent experience, in a STEM field such as science, technology, engineering, mathematics, or physics.
- At least 5 years of proven experience, or a master's degree.
- Experience with operating-system and server-level automation, CI/CD processes, and DevOps using Python, shell scripting, Ansible, Jenkins, C/C++, Java, and JavaScript.
- Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
- Hands-on experience with model testing, AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, as well as NLP and LLM benchmarking.
- Experience using AI development tools to create test plans, develop test cases, and automate test cases.
- Background in firmware, BMC/OpenBMC, network protocols, enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, and UEFI.
- Experience with Redfish, GitHub, GitLab, Gerrit, PXE, SLURM, Docker, and Kubernetes is a plus.
Preferred Qualifications
- Experience with AI-related tools, LLMs, and NLP.
- Experience working with NVIDIA GPU hardware.
- Understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
- Background in parallel programming, ideally CUDA or OpenCL.
Compensation and Benefits
The base salary range is $140,000–$224,250 USD for Level 3 and $168,000–$270,250 USD for Level 4. Compensation is determined by location, experience, and pay for employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 24, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
NVIDIA Spring 2027 Internships: Developer and Performance Technology
Nvidia · Santa Clara, United States
USD 20-71 per hour
NVIDIA 2027 Internships: Software Engineering
Nvidia · Santa Clara, United States
USD 20-71 per hour
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year