Solutions Architect - Rack Scale AI Systems

at Nvidia
USD 208,000-414,000 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Android Communication @ 6 Debugging @ 7 Deep Learning GPU @ 4 InfiniBand @ 4 Linux @ 4 MPI @ 4 Networking @ 4 System Architecture @ 4

Details

NVIDIA's Infrastructure, Planning and Process (IPP) Cloud Infrastructure Team is seeking a Solutions Architect to design and deploy infrastructure solutions for Rack Scale AI products. IPP is a global organization supporting NVIDIA's Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence, and Driverless Cars teams. Its cloud services process nearly half a million automated jobs per day across thousands of servers and support heterogeneous environments with Windows, Linux, and Android operating systems, NVIDIA GPUs, and Tegra processors.

Responsibilities

  • Work with NVIDIA product teams to understand new product roadmaps and requirements, primarily for Rack Scale AI products.
  • Identify optimum solutions for deploying products in data center and laboratory environments using sophisticated design techniques, services, and tools.
  • Assist with rolling out and deploying development features that support the latest NVIDIA hardware and technologies.
  • Collaborate with engineers, architects, technical product managers, and application developers to establish strategies for product launches.
  • Define and implement full-scale solutions for product onboarding into hosted and private cloud environments.
  • Solve complex problems involving multi-site deployments of NVIDIA products and improve deployment quality and time to market.
  • Collaborate with system engineering, software engineering, mechanical and thermal engineering, operations, data center teams, external vendors, and other partners to deliver reliable platforms from concept through prototype and deployment.
  • Integrate and optimize cluster deployment methods and manage software stack deployments, including provisioning services into the cloud.

Requirements

  • Bachelor's or Master's degree in Computer Science, Software Engineering, or equivalent experience.
  • 12 or more years of relevant experience.
  • 6 or more years of Linux and scripting experience.
  • Solid background in operating system kernels and system engineering.
  • Ability to quickly understand technologies outside one's domain and deploy systems in complex hardware-to-software configurations in a fast-paced environment.
  • Strong technical understanding of embedded systems, orchestration and automation systems, data centers, and cloud architecture.
  • Excellent communication and planning skills.
  • Strong problem-solving ability and experience in product engineering, failure analysis, debugging, hardware or test design.
  • Understanding of dense data center design, including compute, storage, and networking.

Additional Qualifications

  • Understanding of software engineering principles and enterprise system architecture.
  • Experience administering and automating GPU and compute clusters.
  • Experience with large-scale quality assurance environments for product bring-up.
  • Experience with large-scale and cluster computing, including MPI.
  • Experience with data center design, high-speed InfiniBand interconnects, cluster storage, and scheduling-related design or management.

Benefits

The role includes eligibility for equity and NVIDIA benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until July 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs