Senior Datacenter Technical Program Manager, At-Scale AI Clusters

at Nvidia
USD 168,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI Deep Learning @ 4 GPU @ 4 Grafana @ 6 HPC Prometheus @ 6 Splunk @ 6

Details

NVIDIA is looking for a highly motivated Technical Program Manager (TPM) to join the Applied Systems Engineering Team and drive datacenter integration for the next generation of NVIDIA AI supercomputing systems. This role will support the lifecycle of AI systems at scale, from datacenter design and requirements definition through systems integration of AI clusters into the datacenter environment and production support.

The role will drive collaboration between engineering leaders across multiple hardware and software teams to build AI supercomputers for NVIDIA engineers and develop reference architectures for customers and partners.

Responsibilities

  • Collaborate with engineers and architects to build and deploy large-scale GPU computing systems based on NVIDIA's reference supercomputing architectures.
  • Lead the integration of new AI clusters with datacenter facilities, including demanding requirements for power, cooling, and instrumentation.
  • Coordinate the design and fit-out of new datacenter builds with internal engineering teams and external contractors.
  • Own and produce detailed documentation for the end-to-end datacenter fit-out and integration process.
  • Communicate with engineering leadership to prioritize and address key issues essential to the success of major customers.

Requirements

  • Bachelor's degree in Applied Science or Engineering, or equivalent experience.
  • 8 or more years of overall experience.
  • Experience with high-performance computing systems and GPU clusters deployed in on-premises datacenters.
  • Passion for understanding challenging technical problems and driving the process of finding solutions.
  • Strong teamwork and interpersonal skills for coordinating workflows across multiple teams.

Preferred Qualifications

  • Understanding of datacenter design, including familiarity with power and cooling technologies.
  • Expertise in system monitoring and instrumentation of large clusters using technologies such as Prometheus, Grafana, Splunk, Modbus, and BACnet.
  • Experience working with engineering or academic research communities supporting high-performance computing or deep learning.

Compensation and Benefits

The base salary range is USD 168,000–258,750 for Level 4 and USD 200,000–322,000 for Level 5. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs