Senior Software Development Engineer in Test

at Nvidia
USD 168,000-270,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 AWS @ 4 Ansible @ 7 Azure @ 4 CI/CD @ 4 Debugging @ 4 Distributed Systems Docker @ 7 Grafana @ 6 HPC HTTP @ 6 Jenkins @ 4 Kubernetes @ 7 Linux @ 6 Networking @ 4 Playwright @ 4 Prometheus @ 6 Security @ 4 Selenium @ 4 Slurm @ 7 Terraform @ 4 Thanos @ 6

Details

NVIDIA is seeking a highly skilled Senior Test Developer / Test Engineer to join its Enterprise Software QA team. The role focuses on the design, construction, optimization, and testing of large-scale infrastructure for NVIDIA unified cloud services and data center offerings. The engineer will work with cloud infrastructure, distributed systems, and AI-assisted testing tools in an innovative environment.

Responsibilities

  • Collaborate with development teams on test plans for all layers of the software stack, including execution, reviews, failure analysis, and assessment of overall quality and risk.
  • Work with customer project managers on software issues and incorporate technical feedback from OEMs and cloud service providers.
  • Develop benchmarks to track execution and implement process improvements to increase efficiency.
  • Use AI development tools to accelerate test scoping, test planning, execution, and automation workflows.
  • Lead NVIDIA cloud and data center bring-up activities, including validation, reporting, issue debugging with engineering teams, design input, and expanding test coverage.
  • Design, develop, and maintain CI/CD pipelines for continuous testing in cloud environments.
  • Perform performance, scalability, and reliability testing of cloud services.
  • Implement and maintain test environments on AWS, Azure, Google Cloud, or OCI Cloud.
  • Monitor infrastructure and configure alerts for significant events to maintain system performance and reliability.
  • Coordinate with partner teams to ensure test cluster availability and lead issue resolution.
  • Ensure the quality of cloud products, with a focus on security, storage, workloads, performance, and the latest software and firmware components.

Requirements

  • Master's or Ph.D. degree in Computer Science or a related field, or equivalent experience.
  • Experience using AI development tools to create and automate test cases, measure code coverage, and perform triage.
  • At least 8 years of hands-on experience with cluster management and related tools, including Docker containers, Slurm, Kubernetes, and Ansible.
  • At least 2 years of experience with cloud infrastructure platforms such as AWS, Azure, Google Cloud, or OCI Cloud.
  • Hands-on experience with networking, storage, security, cluster configuration and debugging, and cloud infrastructure management tools such as Terraform and Ansible.
  • Expertise administering, operating, and configuring Kubernetes.
  • Experience with CI/CD tools such as GitLab and Jenkins, as well as the GitOps model.
  • Proficiency with monitoring tools including Prometheus, Grafana, CloudWatch, and Thanos.
  • Proficiency debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.

Additional Qualifications

  • Familiarity with Base Command Manager for managing and monitoring high-performance computing environments.
  • Experience writing web application automation using tools such as Selenium or Playwright.

Compensation and Benefits

The base salary range is USD 168,000–270,250 per year, determined by location, experience, and compensation for similar positions. The role also includes eligibility for equity and benefits.

More jobs at Nvidia

Similar jobs