Engineering Manager, Reliability Engineering Flywheel - EDA Infrastructure

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 6 LLM Leadership @ 6 Security Technical Leadership

Details

NVIDIA’s EDA Infrastructure organization builds and operates the systems that support chip development. The team is responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness.

The Engineering Manager will own the roadmap and delivery, from defining how teams work to building the tools they use. The role partners with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. The team uses automation, AI, and lessons from operational events to drive improvements.

Responsibilities

  • Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results.
  • Set technical direction, prioritize work, and guide execution across engineering and operational disciplines.
  • Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness.
  • Hire and develop engineers and technical leads, building a team with clear ownership and accountability.
  • Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents.

Requirements

  • BS degree or equivalent experience with 10+ overall years of software engineering or related experience, including 5+ years of engineering leadership managing teams or complex technical programs.
  • Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement.
  • Strong technical judgment in software architecture, platform integration, and engineering tradeoffs.
  • Clear communication with engineers, cross-functional partners, and executive stakeholders.
  • A record of developing engineers, growing teams, and delivering results under pressure.

Preferred Qualifications

  • Established readiness standards covering service ownership, support coverage, and reliability objectives.
  • Experience building, integrating, and scaling platforms for incident management, maintenance, customer experience management, on-call, and production readiness.
  • Experience applying AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation.
  • Experience supporting EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements.

Benefits

  • Equity and benefits are provided.
  • Applications will be accepted at least until September 26, 2026.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs