Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
Communication @ 6
Compliance
GPU @ 4
Jira @ 6
Kubernetes @ 4
Machine Learning
Observability
Security
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an accomplished Technical Program Manager to join the DGX Cloud team. The role focuses on cloud infrastructure bring-up with external partners and building AI capacity and infrastructure globally in collaboration with cloud service providers, NVIDIA Cloud Providers, and internal engineering teams.
Responsibilities
- Partner with storage and network engineering teams to define and communicate requirements to Cloud Service Providers and NVIDIA Cloud Providers.
- Drive alignment and plans of record for capacity blocks based on workload needs.
- Engage early with cloud providers to understand managed storage and network solutions and align them with the NVIDIA Cloud roadmap.
- Gather technical requirements, develop roadmaps, establish milestones, and ensure adherence to the Product Lifecycle process.
- Manage ongoing capacity operations and engineering engagement with cloud-provider partners, focusing on availability, maintenance, and critical performance indicators.
- Understand workload requirements and related hardware and infrastructure needs, including speeds and feeds required to optimize infrastructure readiness with cloud vendors and NVIDIA Cloud Providers.
- Use Jira and other program management platforms to provide rigor and structure for engineering deliverables.
- Drive adoption of third-party and in-house cloud infrastructure solutions for deployment, support, security, compliance, and observability across DGX Cloud.
- Establish key performance indicators and quantitatively demonstrate program value and impact.
- Identify, resolve, and mitigate risks and issues affecting scope, schedule, and quality.
- Promote continuous improvement and identify process improvements within cloud infrastructure operations.
Requirements
- 10+ years of technical program management experience, including planning and executing large-scale cloud infrastructure programs with external organizations.
- Strong focus on software engineering projects within a matrixed organization.
- Extensive hands-on cloud infrastructure experience, preferably at a major Cloud Service Provider.
- Domain knowledge of the bring-up and end-to-end operation of compute, storage, and GPU infrastructure, including common hardware and software failure points.
- Expert-level proficiency with Jira, Smartsheet, or similar program management tools.
- Strong strategic and tactical thinking, consensus-building, and program execution skills.
- Ability to work effectively in ambiguous environments.
- Excellent communication and technical presentation skills, including for executive audiences.
- BS or MS in Electrical Engineering or Computer Science, or equivalent experience.
Preferred Qualifications
- In-depth knowledge of NVIDIA GPU products, including deployment and bring-up.
- Working knowledge of Kubernetes, API integration, Terraform, and other cloud technologies.
- Experience identifying and addressing process improvement opportunities.
- Significant experience with productivity tools and process automation.
- Familiarity with cloud-native product and services environments, AI, ML infrastructure, and cloud services.
Benefits
- Base salary range of $168,000–$258,750 for Level 4 or $200,000–$322,000 for Level 5, depending on location, experience, and comparable employee compensation.
- Eligibility for equity and benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Architect, Networking AI
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Manager, CUDA Driver
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Custom SoC IP Verification Engineer
Nvidia · Santa Clara, United States
USD 168,000-310,500 per year
Research Intern, Fundamental Generative AI - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Partner Solutions Architect
Nebius · Canada, United States
USD 250,000-320,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff - Cloud Infrastructure
SpaceXAI · Washington, United States, Palo Alto, United States
USD 180,000-440,000 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 156,000-215,000 per year
Senior AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 132,000-182,000 per year