Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Debugging @ 4
Linux @ 3
MLOps @ 4
Machine Learning
Networking @ 7
Python @ 3
QA @ 4
Stress Testing
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Systems Software Test Engineer to join the Cloud Service Provider (CSP) Engagements team. The role focuses on ML software stack validation for datacenter products such as GB200 and Vera Rubin. It combines full-stack validation from cluster to rack scale with customer-facing responsibilities, enabling cloud service providers to deploy high-performance training and inference platforms. You will validate stable and performant technical solutions from concept through deployment, working across hardware and software layers.
Responsibilities
- Define test strategies and validation plans for CSP integration milestones.
- Partner with hyperscalers to understand their testing methodologies, identify gaps, and provide NVIDIA recommendations.
- Reproduce, characterize, and triage customer bugs in customer environments.
- Review internal test plans and results, and publish summary test reports for rack-scale product releases during NPI phases.
- Validate fixes, mitigations, and release updates against deployed CSP software modules and known-good partner configurations.
- Partner with NVIDIA development teams on root-cause analysis and confirm release readiness with clear pass/fail evidence.
- Collaborate with CSP teams on provisioning, access, break-fix workflows, and environment readiness.
- Produce release-readiness summaries for internal stakeholders and partner-facing engineering reviews.
- Manage large testing-output datasets and develop tooling for efficient debug-data retrieval, visualization, and reporting.
- Localize problems with customers using targeted reproduction steps, stress testing, and edge-case testing.
- Run performance benchmarks for training and inference workloads.
- Collaborate with AE, FAE, and Solution Architect teams on customer-issue validation and technical documentation.
- Replicate reported problems in the local lab.
Requirements
- Experience in validation, QA, system testing, diagnostics, platform bring-up, or release qualification for complex hardware-software systems.
- Strong understanding of server platforms, firmware, drivers, operating-system integration, networking, and large-scale cluster environments.
- Hands-on experience debugging issues across hardware, firmware, software, networking, and infrastructure layers.
- Ability to analyze logs, telemetry, diagnostic outputs, automation failures, and system health signals.
- Familiarity with Linux environments, shell scripting, Python or similar automation, and CI/regression workflows.
- Experience creating test plans, regression suites, validation reports, and defect documentation.
- Strong cross-functional communication skills with QA, development, field, support, and customer engineering teams.
- Proficiency in Python, with a strong background in test automation and test infrastructure design.
- Ability to communicate and collaborate effectively with partner and customer teams.
- Bachelor’s or master’s degree in Computer Engineering, Computer Science, or a related field, or equivalent experience.
- 8+ years of system software validation experience.
Preferred Qualifications
- Hands-on experience with cloud and cluster-level deployment and MLOps.
- Experience running deep-learning workloads and related automation.
Benefits
- Equity and benefits are included.
- Applications will be accepted at least until August 3, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior QA Automation Engineer, Network Simulation Platform
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Platform Security Engineer, OpenBMC
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year