Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 6
Customer Support @ 3
GPU @ 4
Leadership @ 4
Machine Learning
Networking @ 8
Salesforce @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Service Technical Program Manager to coordinate large-scale service operations programs for NVIDIA GPU and networking platforms across global server facilities. The role supports production deployments, ongoing operations, retrofit and upgrade programs, and serviceability improvements for hyperscale customers. It bridges product engineering, field service, supply chain, and customer operations to ensure NVIDIA systems remain performant, serviceable, and continuously optimized in live production environments.
Responsibilities
- Lead end-to-end execution of large-scale service programs involving GPU-driven compute platforms and advanced networking technologies across production data centers.
- Support in-field service operations, including incident management, blocking issues, warranty delivery, and scalable hardware replacement strategies for field-replaceable units.
- Plan and coordinate infrastructure upgrades and retrofit programs, including system refreshes, field hardware swaps, firmware rollouts, and architecture changes.
- Collaborate with hyperscale customers on service strategies for rack-level deployment, cluster expansion, and failure-domain development.
- Collaborate with NPI, Product Engineering, Hardware, and Networking teams to embed serviceability principles into products.
- Support NPI-to-production transition planning, including spares, repair workflows, tooling, and global logistics.
- Serve as a primary point of contact for hyperscale customers when applicable, handling field-blocking issues and driving continuous service improvements.
- Apply data and operational insights to improve efficiency and service delivery across deployed systems.
- Enable global field and regional teams through runbooks, training, tools, and standardized processes for compute and network infrastructure support.
- Build customer relationships through onsite and remote engagement, improve service delivery, optimize retrofit programs, and drive measurable operational success.
Requirements
- Bachelor's degree or equivalent experience in Electrical, Mechanical, Computer, or Network Engineering, combined with strong program management expertise in planning, execution, risk management, and delivery of complex technical programs across a global customer base.
- 15+ years of experience supporting hyperscale or cloud data center environments, including compute infrastructure and networking technologies.
- Practical experience leading production data center activities such as hardware installation, rack setup, cluster initialization, and lifecycle oversight.
- Experience supporting GPU-based compute platforms and/or large-scale networking infrastructure in live environments.
- Experience in field service operations, including incident management, FRU/CRU hardware replacement methods, and large-scale service delivery.
- Experience leading retrofit, upgrade, and refresh programs across distributed infrastructure fleets.
- Ownership of end-to-end service programs spanning NPI readiness, deployment, sustaining operations, and end-of-life transitions.
- Understanding of service logistics, spares planning, and global supply chain coordination.
- Familiarity with customer support systems such as Salesforce/SFDC and data-driven tools for operational insights and decision-making.
Preferred Qualifications
- Experience working directly within hyperscale providers and supporting compute and networking infrastructure at scale.
- Ability to apply AI/ML-based tools to improve program clarity, optimize operations, and enable predictive decision-making.
- Leadership experience involving engineering, field operations, distribution, and customer-facing teams.
- Ability to translate operational challenges into product improvements and serviceability enhancements.
- Analytical approach to service optimization and operational excellence.
- Clear communication skills and the ability to influence technical teams and executive collaborators.
Benefits
- Base salary range: $240,000–$379,500 USD, determined by location, experience, and pay for employees in similar positions.
- Eligible for equity and benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Principal Systems Software Engineer, Semiconductor Systems Inspection
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Architect, Agentic AI for Marketing
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Software Engineer
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Software Engineer, Performance Lab and Tools
Nvidia · Santa Clara, United States
USD 116,000-189,800 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Similar jobs
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Principal System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Principal High-Performance LLM Training Engineer
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Platform Security Engineer, OpenBMC
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-405,000 per year
Technical Program Manager, Silicon
Anthropic · New York City, United States, San Francisco, United States
USD 365,000-435,000 per year