Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CI/CD @ 7
Communication @ 6
GPU @ 7
Jira @ 6
Kubernetes @ 7
Leadership @ 6
Networking
Observability @ 4
QA @ 6
Reporting @ 6
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an experienced and talented Technical Program Manager for NVIDIA's DGX Cloud to help deliver value to DGX Cloud customers.
As a DGX Cloud Technical Program Manager, you’ll be a key partner to Engineering, Infrastructure, and Software teams, driving critical cloud infrastructure programs across DGX Cloud. You’ll mature how we bring up AI capacity—strengthening process resiliency, driving automation into our PLC and acceptance workflows, and enabling early access to upcoming NVIDIA platforms.
Responsibilities
- Lead the end-to-end execution of NPI programs across engineering, operations, and cloud service provider (CSP) partners
- Lead the DGX Cloud NPI Early Access Program—enabling processes that lead to engineering teams across DGXC getting early access to critical NVIDIA systems (VR / VR Ultra) in order to develop critical software and automation
- Drive PLC process for capacity bring-up into system-based, automated solutions—taking a baseline PLC type process and driving iterative improvements and system approaches to codifying in tooling including Jira
- Coordinate site readiness and infrastructure bring-up activities, including networking, inventory, corporate IT, and security integration
- Partner with SW stack teams to track development, testing, and integration across product phases
- Define and implement acceptance testing, validation workflows, and readiness gates for new platforms
- Work closely with stakeholders to develop scalable NPI processes, tools, and dashboards
- Drive automation efforts for break/fix workflows, telemetry enablement, and system health validation
- Facilitate regular communication with leadership, engineering, CSP teams, and Colo partners and cultivate a culture of continuous improvement and process innovation
Requirements
- 12+ years of technical program management experience, with a focus on infrastructure, hardware/software integration, or cloud platforms
- Success in leading NPI or large cross-functional programs in fast-paced environments
- Experience working with cloud service providers, large-scale data center deployments, or enterprise-scale infrastructure programs
- Strong understanding of GPU compute, Kubernetes, CI/CD pipelines, and cloud-native services
- Demonstrated experience building or improving product development processes and team workflows
- Skilled in tools such as Jira, Confluence, dashboards, and reporting tools
- Ability to influence cross-functional teams, including HW, SW, QA, Site Ops, and Product
- Outstanding communication and leadership skills, capable of collaborating effectively with senior collaborators
- BS/MS in CS, EE, related technical field, or equivalent experience
Ways to stand out from the crowd
- Experience in launching cloud infrastructure products or large-scale hardware-software systems
- Previous involvement in New Product Introduction (NPI), including platform bring-up and validation
- Familiarity with AI infrastructure, or GPU-based cloud platforms
- Experience with process automation, observability (telemetry/metrics), and health check frameworks
- Passion for building repeatable systems, tools, and cross-org efficiency at scale
Additional information
- Base salary range: 168,000 USD - 258,750 USD for Level 4; 200,000 USD - 322,000 USD for Level 5
- Equity and benefits are also included
- Applications accepted at least until July 26, 2026
- This posting is for an existing vacancy
More jobs at Nvidia
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Director, Autonomous Vehicles Platform
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Developer Technology Engineer - Edge Agentic Ai
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - Halos
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Drive Os Communication Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Principal Security Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 347,000-490,000 per year
Staff Forward Deployed Engineer
GitLab · United States
USD 254,000-297,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Software Engineer, Attestation Services - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 347,000-490,000 per year
Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Security Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Software Engineer, Compute Infrastructure
OpenAI · New York City, United States, San Francisco, United States, Seattle, United States, United Kingdom
USD 230,000-405,000 per year