Senior Manager, Performance Engineering – Kernel and Software Platforms

at Nvidia
USD 272,000-488,800 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 6 API @ 6 CUDA @ 4 Codex @ 6 GPU GenAI Generative AI @ 6 LLM @ 4 LLVM @ 4 PyTorch Python @ 6

Details

NVIDIA’s accelerated computing platform relies on continuous performance excellence throughout development. This role leads an engineering team responsible for supervising and optimizing product performance across the full hardware lifecycle, from early simulation and emulation through post-silicon bring-up, production hardware, and ongoing release support. The role analyzes kernel authoring flows from domain-specific languages to internal code representations, establishes performance expectation models, curates workload testlists, coordinates with CUDA release schedules, and promotes automation using AI tools such as Claude and Codex.

Responsibilities

  • Lead end-to-end performance tracking for GPU software products across pre-silicon builds, simulation, emulation, initial post-silicon validation, product hardware, and post-release maintenance.
  • Architect theoretical and empirical performance models to establish early design targets.
  • Correlate pre-silicon simulation predictions with early hardware and production silicon to diagnose and eliminate performance discrepancies.
  • Evaluate and benchmark performance translation across kernel authoring flows, including Triton and PyTorch, compiler intermediate representations, and target hardware execution.
  • Identify, create, and maintain stress-test suites and workload testlists representative of production applications.
  • Use workload testlists to detect performance regressions early in simulation and validate hardware release candidates.
  • Collaborate with compiler, architecture, and platform software teams to align performance delivery with CUDA release schedules.
  • Integrate AI infrastructure, including Claude, OpenAI Codex, and agentic LLM workflows, to automate telemetry analysis, investigate pre- versus post-silicon performance differences, and streamline reporting pipelines.

Requirements

  • Master’s or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • 10 or more years of experience in systems or software performance engineering, platform benchmarking, or a related area.
  • 5 or more years of experience leading or managing technical engineering teams.
  • Demonstrated experience tracking performance across the full hardware pipeline, from design-time simulation and emulation infrastructure through post-silicon bring-up and deployed production hardware.
  • Experience building performance expectation models and correlating simulation predictions with physical hardware telemetry.
  • Understanding of modern kernel compilation pipelines and compiler flows from DSL to IR to target code, including the impact of high-level software abstractions on low-level execution efficiency.
  • Experience developing workload testlists to detect performance regressions and aligning performance delivery with major software release cycles such as CUDA.
  • Proficiency in Python automation and practical experience using generative AI APIs or models, including Codex, Claude, or custom agents, to automate triage and analytical workflows.

Preferred Qualifications

  • Experience building automated shift-left performance validation frameworks that map pre-silicon simulator data to post-silicon measurements.
  • Hands-on analytical experience with MLIR, LLVM IR, NVVM, or PTX.
  • Experience designing LLM-driven agents that analyze performance regressions between hardware and software releases and summarize root causes.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs