Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Compliance @ 4
Distributed Systems @ 4
GPU @ 4
Grafana @ 4
Leadership @ 4
Machine Learning
Networking @ 4
Observability @ 4
Prometheus @ 4
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
For over 25 years, NVIDIA has led the world in visual computing and accelerated computing. The DGX Cloud organization builds and operates the AI infrastructure that supports generative models, autonomous systems, and large-scale research. NVIDIA is seeking a Technical Program Management Manager to lead core infrastructure programs across DGX Cloud, including network, storage, trust services, security, break/fix operations, and telemetry. This role manages a team of Technical Program Managers responsible for bringing structure, operational rigor, and cross-functional alignment to infrastructure programs that keep DGX Cloud resilient, scalable, and customer-ready.
Responsibilities
- Lead and nurture a team of Technical Program Managers engaged in DGX Cloud core infrastructure projects.
- Drive progress across network, storage, trust services, security, telemetry, and break/fix operational workstreams.
- Partner with engineering, product, operations, security, and cloud provider teams to define priorities, achievements, dependencies, and delivery plans.
- Build operating rhythms for infrastructure planning, blocking-issue management, risk tracking, and cross-functional decision-making.
- Improve access to infrastructure health, delivery status, blockers, and program risks through practical metrics, dashboards, and reporting.
- Coordinate break/fix and operational readiness programs to improve reliability, response time, and customer impact management.
- Support continuous improvement across TPM practices, helping standardize planning, execution, and communication across DGX Cloud infrastructure.
Requirements
- More than 12 years of experience in technical program management, infrastructure program management, or similar roles, including more than 3 years directing or supervising TPMs.
- Experience managing infrastructure programs in networking, storage, security, trust services, observability, telemetry, or cloud operations.
- Ability to manage priorities, dependencies, risks, and execution plans across multiple engineering teams.
- Experience building TPM operating rhythms, including status reviews, blocking-issue processes, achievement tracking, and leadership-ready updates.
- Working knowledge of cloud infrastructure, distributed systems, or large-scale platform operations.
- Strong communication skills and the ability to translate complex infrastructure work into clear program status, risks, and decisions.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent experience.
Preferred Qualifications
- Experience supporting infrastructure for AI/ML platforms, GPU clusters, or large-scale cloud services.
- Experience with observability and telemetry tools such as Grafana, Prometheus, or similar platforms.
- Experience with security, trust, compliance, or reliability programs in cloud infrastructure environments.
- Track record of improving operational processes for break/fix, incident response, or infrastructure readiness.
- Strong technical judgment and the ability to partner closely with engineering leaders while developing TPM talent.
Benefits
- Base salary range: USD 240,000–379,500 per year, determined by location, experience, and the pay of employees in similar positions.
- Eligibility for equity and benefits.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year