Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Leadership @ 7
Networking @ 7
Project Management @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's AI Factory Infrastructure team develops global reference designs, de-risks technologies needed for the next generation of compute and network products, and builds AI infrastructure at scale to validate solutions. This role leads highly cross-functional teams spanning hardware, software, and facility infrastructure to deliver solutions and infrastructure.
Responsibilities
- Collaborate with product owners and technical leads to identify and collect requirements for next-generation AI factories.
- Build, supervise, and complete long-term programs, including schedules, resourcing, and checkpoints.
- Work with data center and hardware teams to find creative solutions to complex problems and co-develop mitigation strategies.
- Lead planning with key internal partners on capacity demands, engineering roadmaps, and data center expansions.
- Own end-to-end delivery of 100MW+ AI factory data center deployments, from construction readiness through commissioning, turn-up, and handoff to operations.
- Coordinate general contractors, colocation providers, utilities, and OEMs to align electrical and mechanical scope, long-lead equipment, and site logistics for large-scale AI cluster deployments.
- Drive integrated readiness reviews and acceptance criteria across power, liquid cooling, networking, and platform/hardware teams to ensure performance and reliability targets are met for AI factory applications.
- Develop program plans for government grants and initiatives.
- Translate program requirements into Basis of Design documents.
- Bring together team members, foster a collaborative approach to delivery, and hold team members accountable for action items and timelines.
Requirements
- Outstanding long-term planning and execution skills for data center lifecycle planning, including large-scale AI factory buildouts and expansions.
- Experience managing end-to-end data center deployments for high-density AI infrastructure, including commissioning, readiness reviews, turn-up, and operational handoff.
- Demonstrated ability to coordinate colocation providers, general contractors, utilities, and OEMs to deliver complex electrical and mechanical scope at 100MW+ campus scale.
- Strong technical and program leadership across power delivery, liquid cooling, networking, and compute/platform teams to define acceptance criteria and ensure performance and reliability targets are met.
- 12+ years of experience providing program and project management leadership for data center projects covering mechanical, electrical, and plumbing construction, with large-scale server, storage, and network deployments.
- BS or MS degree in Engineering, or equivalent experience.
Preferred Qualifications
- In-depth knowledge of infrastructure hardware and software, as well as electrical and mechanical data center facility technologies.
- Familiarity with NVIDIA's AI compute technology stack and the ability to translate platform requirements into data center infrastructure designs covering power delivery, liquid cooling, space, and network topology at scale.
- Experience with colocation data center environments.
Compensation and Benefits
The base salary range is USD 200,000–322,000 for Level 5 and USD 240,000–379,500 for Level 6. The position also includes eligibility for equity and benefits. Applications will be accepted at least until September 24, 2026.
More jobs at Nvidia
Senior Software Engineer, Capacity Management - DGX Cloud
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Poland
PLN 221,200-507,000 per year
Tech Lead - Cryptographic Asset Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NVIS Strategy Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Compiler Engineer, Agentic Compiler Systems
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Technical Program Manager
Nebius · United States
USD 115,000-275,000 per year
Technical Program Manager – Chip System Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Principal Architect, System Software - Orbital Data Center
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Data Center Manager
Nebius · Philadelphia, United States
USD 115,000-275,000 per year
Technical Program Manager, Silicon
Anthropic · New York City, United States, San Francisco, United States
USD 365,000-435,000 per year
Senior Executive IT Support Engineer, Tech Foundations
Airbnb · San Francisco, United States
USD 132,000-155,000 per year
Senior Software Embedded Engineer - Networking
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior DevOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year