Head of Global Customer Engineering - Internal EDA Infrastructure
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Communication @ 9
GPU
GenAI
Generative AI @ 4
HPC
Kubernetes @ 4
LLM @ 3
Leadership @ 9
Machine Learning
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Every NVIDIA GPU starts as a job on NVIDIA's internal EDA infrastructure. Thousands of hardware engineers designing next-generation silicon depend on this infrastructure being fast, available, and frictionless. NVIDIA is hiring a hands-on engineering leader to develop Global Customer Engineering for Internal EDA Infrastructure, the organization positioned between chip designers and the platforms they use.
This multi-team leadership role spans managers and technical leads across a globally distributed organization and has two co-equal mandates: building a world-class customer engineering organization for EDA infrastructure and leading the Kubernetes platform team that delivers one of its customer services end to end.
The role involves designing software solutions and automation pipelines to eliminate operational overhead. The leader is expected to remain connected to live production environments and engineer permanent automated remedies for recurring customer engineering toil.
Responsibilities
- Build and develop the leaders managing a globally distributed team of customer-facing engineers responsible for the reactive support queue for EDA infrastructure.
- Ensure escalations into service engineering contain clean, high-quality signal.
- Lead the Kubernetes platform team as a first-class customer engineering service, owning availability, roadmap, developer experience, and operational excellence.
- Own the evolution of the L0 AI support agent from an FAQ bot into a system that safely resolves real issues against internal systems and measurably reduces human support queue volume.
- Define and defend service-level objectives for operational performance and secure stakeholder buy-in.
- Drive deep support and continuous improvement across EDA job schedulers, including LSF and Slurm, Kubernetes platforms, and internal LLM harnesses used by chip designers.
- Influence technical roadmaps by converting support trends into engineering requirements that ship.
- Use automation and AI/ML to improve support efficiency and reduce ticket volume.
Requirements
- 12+ years of experience building and leading engineering support organizations spanning reactive and proactive customer engineering.
- 5+ years of experience leading a globally distributed organization, including developing managers and leaders.
- Deep technical familiarity with Kubernetes, EDA workloads, job schedulers, and AI/LLM infrastructure foundations.
- Exceptional communication skills, with credibility among individual engineers and executive leadership.
- Bachelor's degree in a relevant engineering field or equivalent experience.
Preferred Qualifications
- Experience leading a platform team and a customer-facing team simultaneously.
- Experience architecting self-healing support ecosystems, including shift-left tooling that resolves issues before tickets are created.
- Experience leading the full lifecycle of generative AI and LLM-based support assistants integrated deeply and safely with internal infrastructure.
- Experience operating LSF, Slurm, or Kubernetes at a scale where manual intervention was a bottleneck, and leading automation to remove that bottleneck.
Company Information
NVIDIA develops technologies in artificial intelligence, high-performance computing, and visualization. The company is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until September 19, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.