Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Ansible @ 7
Azure @ 4
CI/CD @ 4
Debugging @ 4
Distributed Systems
Docker @ 7
Grafana @ 6
HPC
HTTP @ 6
Jenkins @ 4
Kubernetes @ 7
Linux @ 6
Networking @ 4
Playwright @ 4
Prometheus @ 6
Security @ 4
Selenium @ 4
Slurm @ 7
Terraform @ 4
Thanos @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a highly skilled Senior Test Developer / Test Engineer to join its Enterprise Software QA team. The role focuses on the design, construction, optimization, and testing of large-scale infrastructure for NVIDIA unified cloud services and data center offerings. The engineer will work with cloud infrastructure, distributed systems, and AI-assisted testing tools in an innovative environment.
Responsibilities
- Collaborate with development teams on test plans for all layers of the software stack, including execution, reviews, failure analysis, and assessment of overall quality and risk.
- Work with customer project managers on software issues and incorporate technical feedback from OEMs and cloud service providers.
- Develop benchmarks to track execution and implement process improvements to increase efficiency.
- Use AI development tools to accelerate test scoping, test planning, execution, and automation workflows.
- Lead NVIDIA cloud and data center bring-up activities, including validation, reporting, issue debugging with engineering teams, design input, and expanding test coverage.
- Design, develop, and maintain CI/CD pipelines for continuous testing in cloud environments.
- Perform performance, scalability, and reliability testing of cloud services.
- Implement and maintain test environments on AWS, Azure, Google Cloud, or OCI Cloud.
- Monitor infrastructure and configure alerts for significant events to maintain system performance and reliability.
- Coordinate with partner teams to ensure test cluster availability and lead issue resolution.
- Ensure the quality of cloud products, with a focus on security, storage, workloads, performance, and the latest software and firmware components.
Requirements
- Master's or Ph.D. degree in Computer Science or a related field, or equivalent experience.
- Experience using AI development tools to create and automate test cases, measure code coverage, and perform triage.
- At least 8 years of hands-on experience with cluster management and related tools, including Docker containers, Slurm, Kubernetes, and Ansible.
- At least 2 years of experience with cloud infrastructure platforms such as AWS, Azure, Google Cloud, or OCI Cloud.
- Hands-on experience with networking, storage, security, cluster configuration and debugging, and cloud infrastructure management tools such as Terraform and Ansible.
- Expertise administering, operating, and configuring Kubernetes.
- Experience with CI/CD tools such as GitLab and Jenkins, as well as the GitOps model.
- Proficiency with monitoring tools including Prometheus, Grafana, CloudWatch, and Thanos.
- Proficiency debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.
Additional Qualifications
- Familiarity with Base Command Manager for managing and monitoring high-performance computing environments.
- Experience writing web application automation using tools such as Selenium or Playwright.
Compensation and Benefits
The base salary range is USD 168,000–270,250 per year, determined by location, experience, and compensation for similar positions. The role also includes eligibility for equity and benefits.
More jobs at Nvidia
Senior Software Solutions Engineer
Nvidia · Poland
PLN 230,200-487,500 per year
Senior Offensive Security Engineer, Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Operations and Administrator
Nvidia · Santa Clara, United States
USD 112,000-218,500 per year
Senior Software Engineer, Agent Simulation and Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Product Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, Golang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year