Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a highly skilled Performance Modeling Architect to lead the architectural definition and improvement of next-generation CPU cache hierarchies and interconnects. The role focuses on creating scalable solutions for automotive and data center systems and building high-fidelity models that govern data movement across silicon, including next-level caches such as L3/System Cache and coherent fabrics.
Responsibilities
- Develop and maintain high-fidelity, cycle-accurate performance models using C++ and SystemC for coherent interconnects and large-scale shared caches.
- Model and analyze performance bottlenecks across small-cluster automotive SoCs and large, multi-mesh data center architectures.
- Evaluate the performance impact of coherency protocols such as CHI, ACE, and proprietary protocols, as well as snooping filters.
- Run and analyze industry-standard benchmarks, including SPEC, MLPerf, and automotive-specific suites, to drive architectural trade-offs.
- Collaborate with build and verification teams to correlate performance models with silicon.
- Work with software teams to optimize drivers for the underlying hardware topology.
Requirements
- Master’s or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or equivalent experience, with a focus on computer architecture and 5+ years of experience.
- Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols.
- Proven experience with C++ or SystemC for cycle-accurate or functional modeling.
- Proficiency in Python or similar scripting languages for processing large datasets, generating performance visualizations, and automating simulation sweeps.
- Understanding of Network-on-Chip topologies, including mesh, ring, and torus, as well as credit-based flow control and arbitration logic.
Preferred Qualifications
- Practical experience managing functional safety requirements under ISO 26262 for automotive chips and power-performance-area limitations of data center hardware.
- Experience defining or using PMU events to debug performance on real silicon or emulators.
- Background in formal verification or mathematical modeling for proving the correctness of complex coherency state machines.
- Experience building internal tools or frameworks to accelerate architectural exploration.
- Knowledge of emerging memory technologies such as CXL and HBM and their interaction with coherent fabrics.
Compensation and Benefits
The base salary depends on location, experience, and the pay of employees in similar positions:
- Level 3: $152,000–$241,500 USD per year
- Level 4: $184,000–$287,500 USD per year
The role is also eligible for equity and benefits. Applications will be accepted at least until May 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior DevOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Director, Enterprise Network Deployment & Automation
Nvidia · Santa Clara, United States
USD 284,000-425,500 per year
Senior System Software Integration and Debug Engineer
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Performance Engineer
Nvidia · Santa Clara, United States
USD 136,000-270,200 per year
AI Compiler Engineer - New College Grad 2027
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Senior Staff Machine Learning Engineer, Ads Ranking
Reddit · United States, Chicago, United States, New York City, United States, Los Angeles, United States, San Francisco, United States
USD 292,500-409,500 per year
IT Operations Engineer, Asset Management
Anthropic · New York City, United States, San Francisco, United States
USD 180,000-230,000 per year
Staff+ Software Engineer, Distributed Systems
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-485,000 per year
Measurement Data Scientist
AppLovin · Palo Alto, United States
USD 98,000-148,000 per year
Staff Machine Learning Engineer (Platform - Identity)
Coinbase · United States
USD 218,000-256,500 per year
PhD Research Intern, Generalist Embodied Agents Research - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Senior Tools Engineer
Nvidia · United States
USD 184,000-287,500 per year
PhD Research Intern, Hardware and Systems Architecture - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour