Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Compliance @ 4
Distributed Systems @ 4
GPU @ 4
Grafana
Leadership @ 4
Machine Learning
Networking @ 4
Observability @ 4
Prometheus
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
For over 25 years, NVIDIA has led the world in visual computing and accelerated computing. Today, we’re crafting the future of AI by driving breakthroughs in generative models, autonomous systems, and large-scale research. The DGX Cloud organization builds and operates the AI infrastructure that makes this innovation possible. We are seeking a Technical Program Management Manager to lead core infrastructure programs across DGX Cloud, including network, storage, trust services, security, break/fix operations, and telemetry. This role manages a team of TPMs responsible for bringing structure, operational rigor, and cross-functional alignment to infrastructure programs that keep DGX Cloud resilient, scalable, and customer-ready.
What you’ll be doing
- Lead and nurture a team of Technical Program Managers engaged in DGX Cloud core infrastructure projects.
- Propel progress across network, storage, trust services, security programs, telemetry, and break/fix operational workstreams.
- Partner with engineering, product, operations, security, and cloud provider teams to define priorities, achievements, dependencies, and delivery plans.
- Build clear operating rhythms for infrastructure planning, managing blocking issues, risk tracking, and cross-functional decision-making.
- Improve access to infrastructure health, delivery status, blockers, and program risks through practical metrics, dashboards, and reporting.
- Coordinate break/fix and operational readiness programs that improve reliability, response time, and customer impact management.
- Support continuous improvement across TPM practices, helping the team standardize planning, execution, and communication across DGX Cloud infrastructure.
What we need to see
- More than 12 overall years in technical program management, infrastructure program management, or similar roles, including upwards of 3 years directing or supervising TPMs.
- Experience managing infrastructure programs in domains such as networking, storage, security, trust services, observability, telemetry, or cloud operations.
- Strong ability to manage priorities, dependencies, risks, and execution plans across multiple engineering teams.
- Experience building TPM operating rhythms, including status reviews, paths for handling blocking issues, tracking critical achievements, and leadership-ready updates.
- Working knowledge of cloud infrastructure, distributed systems, or large-scale platform operations.
- Strong communication skills with the ability to translate complex infrastructure work into clear program status, risks, and decisions.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field, or equivalent experience.
Ways to stand out from the crowd
- Experience supporting infrastructure for AI/ML platforms, GPU clusters, or large-scale cloud services.
- Background with observability and telemetry tools such as Grafana, Prometheus, or similar platforms.
- Experience with security, trust, compliance, or reliability programs in cloud infrastructure environments.
- Track record improving operational processes for break/fix, incident response, or infrastructure readiness.
- Strong technical judgment and ability to partner closely with engineering leaders while developing TPM talent.
Additional information
- #Li-Hybrid
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 240,000 USD - 379,500 USD.
- You will also be eligible for equity and benefits.
- Applications for this job will be accepted at least until July 25, 2026.
More jobs at Nvidia
Senior System Software Engineer, Interactive World Models
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer – Streaming
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Software Engineer - GeForce Now Low Latency Streaming
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Gpu Architecture Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Compiler Engineer, Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Manager, Software Engineering - Production AI Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Staff Full Stack Engineer, Identity
Stripe · South San Francisco, United States, New York City, United States, Seattle, United States
USD 224,000-336,000 per year
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Perplexity AI · New York City, United States, Ireland, London, United Kingdom, San Francisco, United States
USD 250,000-485,000 per year
Staff Forward Deployed Engineer
GitLab · United States
USD 254,000-297,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year