Distinguished Engineer, End-to-End Scaling Performance Architecture
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 7
Communication @ 7
GPU
HPC
Leadership @ 7
Mentoring @ 7
NVLink @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Distinguished Engineer to join its architecture organization and help define how future accelerated computing systems scale from a single processor to multi-die, multi-GPU, and multi-node platforms. The role sets long-term performance strategy across applications, systems, and architecture, including DRAM, NVLink, and chip-to-chip (C2C) interconnects.
The focus is on architectural direction and application outcomes. The successful candidate will identify how data movement, communication, memory behavior, topology, and compute limit scaling, and translate those insights into priorities that guide multiple product generations. Domain teams own detailed implementation and delivery, while this role aligns their decisions around a shared end-to-end strategy.
Responsibilities
- Define the multi-generation strategy for application scaling across DRAM, NVLink, C2C, compute, and the supporting software stack.
- Translate the behavior of important AI, HPC, and accelerated computing applications into architectural requirements, performance targets, and investment priorities.
- Analyze how bottlenecks shift as workloads scale across dies, GPUs, nodes, model sizes, data sets, and communication patterns.
- Evaluate system-level trade-offs across bandwidth, latency, capacity, topology, coherence, power, area, cost, programmability, and resiliency.
- Establish common workload scenarios, scaling metrics, models, and decision frameworks so architecture teams can compare proposals against application outcomes.
- Identify architectural discontinuities and emerging technology opportunities early enough to shape product and technology decisions.
- Align DRAM, NVLink, C2C, GPU, CPU, system, and software architects around shared performance limits and high-value opportunities.
- Collaborate with application, framework, compiler, runtime, modeling, and post-silicon teams to connect measured behavior with future architecture choices.
- Provide recommendations to senior technical and business leaders, including assumptions, sensitivities, risks, and expected impact.
- Mentor system performance architects, strengthen technical communities across teams, generate sustained intellectual property, and help influence the direction of large-scale accelerated computing.
Requirements
- MSEE, MSCE, PhD, or equivalent experience in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
- 18+ years of relevant industry or academic experience, including experience setting architecture direction for complex, high-performance systems.
- Deep understanding of system performance and scaling, including interactions among DRAM behavior, high-bandwidth fabrics such as NVLink, and C2C communication.
- Strong application-level intuition and the ability to connect workload algorithms, parallelism, communication, locality, and data movement to architecture choices and measurable outcomes.
- Experience with workload characterization, analytical or simulation-based performance modeling, bottleneck analysis, and architecture trade-off evaluation.
- A record of identifying cross-domain opportunities that may not be visible when teams optimize individual components separately.
- Demonstrated ability to create and advance a multi-generation technical strategy through influence across silicon, systems, software, and application teams.
- Clear communication and sound judgment in ambiguous technical areas, with the ability to explain complex system trade-offs to specialists and executive leaders.
- Experience mentoring senior engineers into broader architecture leadership roles and building strong technical communities.
Preferred Qualifications
- Experience shaping product or technology roadmaps around application-level scaling needs.
- Experience helping architecture teams build a quantitative understanding of bottlenecks across memory, interconnect, compute, and software.
- Experience identifying high-impact trade-offs early enough to guide product, architecture, or technology investment.
- Examples of delivering measurable end-to-end improvements in performance, efficiency, or scaling for priority applications.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until August 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.