Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Agile
Communication @ 9
GPU @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Hardware Infrastructure is seeking a Technical Program Manager to lead Infrastructure Capacity Management programs and workstreams. Given this Infrastructure directly supports our near-term and long-term chip roadmap, it must be highly reliable, performant and efficient for our internal users. This is a fast paced and evolving landscape that requires a TPM to guide engineering roadmaps to be delivered with high quality outcomes and a strong foundation of operational excellence. They will partner both internally within Hardware Infrastructure and externally with senior management and our HW Engineering partners to manage capacity operations and scale processes supporting our next stage of growth. They will develop and standardize planning, reporting and execution methodologies and metrics to enable meeting the challenging objectives.
Responsibilities
- Own end-to-end capacity management strategy and execution for EDA Farm, including server procurement, vendor negotiations, server capacity allocation, data center space & power, delivering measurable efficiency and cost optimization across the organization
- Identify and help drive implementation of improvements to EDA Farm infrastructure tooling, automation, and workflows that accelerate server provisioning, reduce manual overhead, and scale capacity management operations
- Drive capacity and procurement initiatives using agile program methodology, aligning planning, prioritization, and delivery across engineering, procurement, and vendor partner teams
- Build and maintain a data-driven capacity model, using metrics and business objectives to improve farm utilization, procurement performance, and vendor SLA consistency — turning insights into actionable cost and capacity optimizations
- Create clear, consistent communication channels that give customers at every level real-time insight into farm capacity health, procurement timelines, supply chain risks, and mitigation plans
- Act as a primary technical and strategic partner between engineering, procurement, finance, and hardware vendors to ensure EDA Farm capacity optimally meets the demands of design and verification teams
Requirements
- B.S. (or equivalent experience) in Electrical Engineering, Computer Science or a related technical field
- 12+ years of proven experience across Capacity Engineering, Capacity Management and/or Technical Program Management roles within the Capacity space
- Experience working with large scale infrastructure with various CPU/GPU architectures both on-prem and cloud
- Exceptional communication and presentation skills for diverse technical and non-technical audiences
- Proactive in identifying and implementing positive changes in both system and process design in a fast-paced environment
Benefits
NVIDIA offers highly competitive salaries and a comprehensive benefits package. You will also be eligible for equity and benefits.