Technical Program Manager, Datacenter

at Nvidia

📍 Santa Clara, United States

USD 132,000-253,000 per year

MIDDLE

✅ On-site

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ^?

Grafana @ 3 Prometheus @ 3 Leadership @ 3 Splunk @ 3 GPU @ 3

Details

NVIDIA is looking for a highly-motivated Technical Program Manager (TPM) to join our Applied Systems Engineering Team to drive datacenter integration for the next generation of NVIDIA AI supercomputing systems. This TPM will play a crucial role throughout the lifecycle of the latest AI systems at scale, from datacenter design and requirements definition, through systems integration of AI clusters into the datacenter environment, and support for these systems as they enter production. This role will drive collaboration between engineering leaders across multiple hardware and software teams, helping us work together to build AI supercomputers for NVIDIA engineers and develop reference architectures to advise customers and partners.

Responsibilities

Collaborate with outstanding engineers and architects to build and deploy large scale GPU computing systems based on NVIDIA's reference supercomputing architectures
Lead the integration of new AI clusters with datacenter facilities with demanding requirements on power, cooling, and instrumentation
Coordinate programs for deploying new cluster architectures, adapting them to changing market requirements, and supporting these architectures as they move into bring-up and production
Document system designs in partnership with multiple engineering groups working on datacenter deployments at scale
Communicate internally with engineering leadership to prioritize and address key issues essential to the success of our largest customers

Requirements

BS (Masters preferred) in Applied Science or Engineering (or equivalent experience)
5+ years of overall experience
Experience with high-performance computing systems and GPU clusters deployed in on-premises datacenters
A passion for understanding challenging technical problems and driving the process of finding a solution
Strong teamwork and interpersonal skills, to facilitate building a collaborative workflow for coordination between many teams

Ways to stand out from the crowd

Understanding of datacenter design, including familiarity with power and cooling technologies
Expertise in system monitoring and instrumentation of large clusters, using technologies such as Prometheus, Grafana, Splunk, Modbus, and BACNet
Experience working with the engineering or academic research community supporting high-performance computing or deep learning

Benefits

The base salary range is 132,000 USD - 253,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis. NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.