Senior Performance Modeling Architect, CPU Fabric and LLC

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI Python @ 6

Details

We are looking for a highly skilled Performance Modeling Architect to lead the architectural definition and improvement of next-generation CPU cache hierarchies and interconnects. The role focuses on creating scalable solutions for automotive and data center systems and building high-fidelity models that govern data movement across silicon, including next-level caches such as L3/System Cache and coherent fabrics.

Responsibilities

  • Develop and maintain high-fidelity, cycle-accurate performance models using C++ and SystemC for coherent interconnects and large-scale shared caches.
  • Model and analyze performance bottlenecks across small-cluster automotive SoCs and large, multi-mesh data center architectures.
  • Evaluate the performance impact of coherency protocols such as CHI, ACE, and proprietary protocols, as well as snooping filters.
  • Run and analyze industry-standard benchmarks, including SPEC, MLPerf, and automotive-specific suites, to drive architectural trade-offs.
  • Collaborate with build and verification teams to correlate performance models with silicon.
  • Work with software teams to optimize drivers for the underlying hardware topology.

Requirements

  • Master’s or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or equivalent experience, with a focus on computer architecture and 5+ years of experience.
  • Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols.
  • Proven experience with C++ or SystemC for cycle-accurate or functional modeling.
  • Proficiency in Python or similar scripting languages for processing large datasets, generating performance visualizations, and automating simulation sweeps.
  • Understanding of Network-on-Chip topologies, including mesh, ring, and torus, as well as credit-based flow control and arbitration logic.

Preferred Qualifications

  • Practical experience managing functional safety requirements under ISO 26262 for automotive chips and power-performance-area limitations of data center hardware.
  • Experience defining or using PMU events to debug performance on real silicon or emulators.
  • Background in formal verification or mathematical modeling for proving the correctness of complex coherency state machines.
  • Experience building internal tools or frameworks to accelerate architectural exploration.
  • Knowledge of emerging memory technologies such as CXL and HBM and their interaction with coherent fabrics.

Compensation and Benefits

The base salary depends on location, experience, and the pay of employees in similar positions:

  • Level 3: $152,000–$241,500 USD per year
  • Level 4: $184,000–$287,500 USD per year

The role is also eligible for equity and benefits. Applications will be accepted at least until May 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs