Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
Communication @ 6
Compliance
GPU @ 4
Jira @ 6
Kubernetes @ 4
Machine Learning
Networking
Observability
Security
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an accomplished and highly skilled Technical Program Manager to join the NVIDIA DGX Cloud team. The role focuses on cloud infrastructure bring-up with external partners and collaboration with Cloud Service Providers (CSPs), emerging cloud providers, and internal engineering teams to build AI capacity and infrastructure globally.
Responsibilities
- Partner with engineering, infrastructure, software teams, and leadership to drive programs related to AI capacity enablement and management.
- Develop and mature foundational capabilities and processes for DGX Cloud, including cluster and capacity bring-up involving CPU, storage, networking, and GPU-supported compute requirements.
- Coordinate with storage and network engineering teams to define and communicate requirements to CSPs and NVIDIA Cloud Providers (NCPs), and drive alignment and plans of record for capacity blocks based on workload needs.
- Engage early with CSPs and NCPs to understand managed storage and network solutions and influence alignment with the NVIDIA Cloud roadmap.
- Gather technical requirements, develop roadmaps, establish milestones, and ensure adherence to the Product Lifecycle (PLC) process.
- Manage ongoing capacity operations and engineering engagement with CSP and NCP partners, focusing on availability, maintenance, and other critical performance indicators.
- Partner internally to understand workload requirements and related hardware and infrastructure needs, including speeds and feeds required to optimize infrastructure readiness with cloud vendors and NVIDIA Cloud Providers.
- Use Jira and other program management platforms to provide rigor and structure for engineering deliverables.
- Drive adoption of third-party and in-house cloud infrastructure solutions for deployment, support, security, compliance, and observability across DGX Cloud.
- Establish key performance indicators (KPIs) and quantitatively demonstrate program value and impact.
- Identify, resolve, and mitigate risks and issues affecting scope, schedule, and quality.
- Promote continuous improvement and identify process improvements within cloud infrastructure operations.
Requirements
- 10+ years of technical program management experience, including planning and executing large-scale cloud infrastructure programs with external organizations.
- Strong focus on software engineering projects within a matrixed organization.
- Extensive hands-on cloud infrastructure experience, preferably gained at a major Cloud Service Provider.
- Domain knowledge of the bring-up and end-to-end operation of compute, storage, and GPUs, including common hardware and software failure points.
- Expert-level proficiency with Jira, Smartsheet, or similar program management tools, with the ability to guide engineering teams in their use.
- Strong strategic and tactical thinking, consensus-building, and program execution skills.
- Ability to work effectively in ambiguous environments.
- Excellent communication and technical presentation skills, particularly with executive audiences.
- BS or MS in Electrical Engineering or Computer Science, or equivalent experience.
Preferred Qualifications
- In-depth knowledge of NVIDIA GPU products, including deployment and bring-up.
- Working knowledge of Kubernetes, API integration, Terraform, and other cloud technologies.
- Experience with productivity tools and process automation.
- Familiarity with cloud-native product and services environments, AI and ML infrastructure, and cloud services.
Benefits
NVIDIA offers attractive compensation, equity, and an extensive benefits package. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment. Applications will be accepted at least until September 5, 2026.
More jobs at Nvidia
Deep Learning Algorithm Engineering Intern - 2026
Nvidia · Zurich, Switzerland
PLN 117,800-204,100 per year
Senior Project Delivery Manager - NVIS
Nvidia · United States
USD 168,000-322,000 per year
Principal Research Scientist, Synthetic Data Generation
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff Software Engineer - Agentic Automation
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Firmware Engineer, NIC Firmware
Nvidia · Seattle, United States
USD 152,000-287,500 per year
Similar jobs
Partner Solutions Architect
Nebius · Canada, United States
USD 250,000-320,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Member of Technical Staff - Cloud Infrastructure
SpaceXAI · Washington, United States, Palo Alto, United States
USD 180,000-440,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year