Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CI/CD @ 4
Communication @ 1
Compliance @ 6
Debugging @ 4
DevOps
GPU @ 6
HPC
Jenkins @ 4
Kubernetes @ 4
Python @ 4
Security @ 6
System Architecture @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Test Architect to join its Enterprise Software QA team and drive the design, construction, optimization, and testing of supercomputers and data center offerings. The role focuses on quality and reliability across NVIDIA's next-generation data center software and firmware stack for AI, HPC, and cloud-scale systems.
Responsibilities
- Define end-to-end test architecture and validation strategies for power features across NVIDIA platforms, from pre-silicon simulation and emulation through post-silicon bring-up and production readiness.
- Develop test plans aligned with product deliverables and customer use cases, influence design decisions to improve testability, and ensure comprehensive test coverage.
- Design and implement modular, reusable test frameworks and automation harnesses for functional, integration, stress, regression, power, security, and performance testing.
- Scale test infrastructure across hundreds of systems in parallel.
- Define quality KPIs, including code coverage, system uptime, bug escape rate, and validation completeness, and establish dashboards and reporting mechanisms.
- Lead root-cause investigations spanning firmware, software, and hardware layers; develop and document debugging methodologies and tools.
- Partner with DevOps and infrastructure teams to improve lab automation and CI/CD pipelines, including nightly and pre-merge continuous testing workflows.
- Validate real-world use cases, customer configurations, and production scenarios, and contribute to release gates and firmware sign-off criteria.
- Mentor software QA engineers and junior test developers while promoting quality, innovation, and continuous learning.
- Evaluate and adopt emerging technologies, embedded validation practices, test automation frameworks, and industry standards.
- Use AI-powered tools and copilots to accelerate test development, automate validation workflows, and streamline debugging and root-cause analysis.
Requirements
- Bachelor's, master's, or Ph.D. degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent experience.
- More than 10 years of experience in data center power enablement related to software or firmware testing, with a focus on telemetry and power efficiency across systems.
- Strong knowledge of system architecture, power shelves, baseboard management, hardware and software power features, industry power standards, system interfaces, and embedded controllers.
- Experience designing test frameworks and infrastructure using Python, C++, or similar languages.
- Expertise with platform standards for security, telemetry, and manageability, including NIST, DMTF, and OCP.
- Hands-on experience with server platforms, networks, storage, cluster configuration, and debugging.
- Background in platform telemetry and data center node lifecycle management and support, including CPU and GPU workloads.
- Proficiency in scripting languages such as Python.
- Experience administering, operating, and configuring Kubernetes and Envoy.
- Experience with CI/CD tools such as GitLab and Jenkins, as well as the GitOps model.
- Experience with lab automation, simulation, hardware-in-the-loop testing, and CI/CD pipelines.
- Strong debugging, problem-solving, and analytical skills.
- Excellent communication and collaboration skills; experience working in globally distributed teams is a plus.
Preferred Qualifications
- Experience with NVIDIA platforms such as DGX, HGX, and Grace Hopper systems.
- Exposure to security validation, FIPS or BMC security compliance, and thermal or power validation.
- Previous experience as a test architect or technical lead for large-scale data center enablement or firmware validation programs.
- Contributions to open-source testing tools or frameworks.
- Strong knowledge of cloud-scale validation, infrastructure automation, or virtualization.
- Experience using AI tools to create agents, design test plans, identify test gaps, automate workflows, and perform failure analysis.
Compensation and Benefits
- Base salary for Level 4: USD 168,000–270,250 per year.
- Base salary for Level 5: USD 200,000–322,000 per year.
- Eligibility for equity and benefits.
- Full-time position.
- NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year