Senior Performance Architect - Heterogeneous Workload Optimization
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
GPU @ 4
Kubernetes @ 4
NVLink
Performance Analysis @ 7
Profiling @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI to define the next era of computing, with GPUs serving as the brains of computers, robots, and self-driving cars.
As EDA workloads transition from traditional CPU-bound tasks to massively parallel GPU-accelerated engines, the complexity of identifying bottlenecks has scaled exponentially. NVIDIA is seeking a Senior Systems Performance Engineer to build next-generation profiling infrastructure. The role involves measuring, analyzing, and optimizing the interaction between extensive design graphs in system memory and high-throughput GPU kernels.
Responsibilities
- Architect and maintain custom profiling frameworks that provide a unified view of execution across CPU environments, including multi-core and multi-socket systems, and GPU environments, including multi-node and NVLink configurations.
- Conduct deep-dive benchmarking of EDA applications to characterize memory access patterns, cache hit rates, and instruction-level parallelism.
- Use GPU profilers to detect inefficiencies such as warp divergence, suboptimal occupancy, and PCIe/NVLink bottlenecks.
- Develop tools to monitor and attribute high-watermark memory usage in multi-terabyte EDA builds, identifying opportunities for data structure compression and smarter memory pooling.
- Develop predictive models to guide hardware procurement and cloud instance selection based on built gate count and algorithmic complexity.
Requirements
- Understanding of the CUDA programming model and experience using GPU profiling tools such as NVIDIA Nsight Systems and NVIDIA Nsight Compute to address PCIe bottlenecks and kernel stalls.
- Extensive knowledge of profiling tools such as perf, eBPF, VTune, or Valgrind, including insight into their internal mechanisms.
- A passion for meticulous benchmarking and the ability to distill sophisticated performance data into actionable engineering roadmaps.
- Experience with distributed compute environments such as Slurm, LSF, or Kubernetes.
- A BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- More than 8 years of relevant experience, including at least 5 years involved in systems-level performance analysis.
Compensation and Benefits
The base salary range is $184,000–$287,500 USD for Level 4 and $224,000–$356,500 USD for Level 5. The role also includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.