Senior Software Development Engineer in Test - Datacenter Server OS
at Nvidia
USD 140,000-270,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Agile
Ansible @ 4
CI/CD @ 7
CUDA @ 6
Debugging @ 7
DevOps @ 7
Docker @ 1
GPU @ 4
GitHub @ 1
Java @ 4
JavaScript @ 4
Jenkins @ 4
Kubernetes @ 1
LLM @ 4
Linux @ 7
Mathematics @ 4
NLP @ 7
OpenCL @ 6
Parallel Programming @ 6
PyTorch @ 4
Python @ 4
Slurm @ 1
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an outstanding individual to join its platform SWQA team. The role requires enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale-out cluster testing, test plan development, AI tools and NLP experience, and DevOps and CI/CD expertise. The candidate should thrive in a diverse work environment, possess strong interpersonal skills, and demonstrate continuous process improvement.
Responsibilities
- Develop and execute NVIDIA HGX, DGX, and MGX platform test plans covering servers, operating systems, firmware, and the CUDA software stack based on design documentation.
- Install and test various operating systems, server firmware, and software stacks.
- Support root cause analysis for reliability and validation test failures, identify root causes, and achieve mitigation.
- Build, develop, and debug server- and OS-level automation frameworks and tests, including front-end and back-end components.
- Review partner and supplier test results and prescribe additional reliability testing for components, servers, and packaging as needed.
- Work in an agile software development team with high production-quality standards.
- Manage the bug lifecycle and collaborate across groups to drive solutions.
Requirements
- Bachelor's degree, or equivalent experience, in a STEM field such as science, technology, engineering, mathematics, or physics.
- Five or more years of proven experience, or a master's degree.
- Experience with OS- and server-level automation, CI/CD processes, and DevOps using Python, Shell, Ansible, Jenkins, C/C++, Java, and JavaScript.
- Strong server and Linux troubleshooting and debugging experience in bare-metal and KVM, VMware, or Hyper-V environments.
- Hands-on experience with model testing and AI tools and frameworks such as TensorFlow, PyTorch, and Cursor, as well as NLP and LLM benchmarking.
- Experience using AI development tools to create test plans, develop test cases, and automate test cases.
- Experience with firmware, BMC/OpenBMC, network protocols, internal and external enterprise storage devices, PCIe buses and devices, I/O sub-devices, CPU and memory, ACPI, UEFI specifications, and Redfish is a major advantage.
- Experience with GitHub, GitLab, Gerrit, PXE, SLURM, Stack, Kubernetes, and Docker is a major advantage.
Preferred Qualifications
- Experience with AI-related tools, LLMs, and NLP.
- Experience working with NVIDIA GPU hardware.
- Solid understanding of Linux virtualization, including KVM and Docker orchestrated with Kubernetes.
- Background in parallel programming, ideally CUDA or OpenCL.
Compensation and Benefits
- Base salary range for Level 3: USD 140,000–224,250 per year.
- Base salary range for Level 4: USD 168,000–270,250 per year.
- Base salary is determined by location, experience, and compensation of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until August 14, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Golang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Deep Learning Frameworks Sustaining Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year