Senior Performance Architect, Nemotron

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Algorithms @ 6 CUDA @ 4 Data Analysis @ 6 Deep Learning @ 4 GPU @ 4 GenAI Generative AI @ 4 LLM @ 4 Machine Learning @ 4 Performance Analysis @ 7 PyTorch @ 4 Python @ 6 Reinforcement Learning @ 4 SGLang @ 4 vLLM @ 4

Details

NVIDIA is looking for a Senior Performance Architect for Nemotron to help shape the next generation of Nemotron models through performance modeling, analysis, and forward projections. The role focuses on AI model–system–hardware co-design and evaluating how architectural choices translate into real-world deployment efficiency. You will help ensure future models achieve Pareto-optimal trade-offs across accuracy, throughput, and interactivity on target platforms.

Recent efforts such as LatentMoE architectures and the Nemotron Super model exemplify the performance-driven co-design this role will advance. The position partners across research, framework development, compiler, and hardware teams to guide decisions affecting production-scale AI efficiency.

Responsibilities

  • Develop high-fidelity analytical performance models to prototype emerging algorithmic techniques and hardware optimizations for the Nemotron family of models.
  • Prioritize features and guide future software and hardware roadmaps based on detailed performance modeling and analysis.
  • Model the end-to-end performance impact of emerging generative AI workflows, including speculative decoding, agentic pipelines, inference-time compute scaling, and reinforcement learning, to understand future data center needs.
  • Keep up with the latest deep learning research and collaborate with deep learning researchers, hardware architects, and software engineers.

Requirements

  • A minimum of a master's degree, or equivalent experience, in computer science, electrical engineering, or a related field.
  • Strong background in computer architecture, roofline modeling, queuing theory, and statistical performance analysis techniques.
  • Solid understanding of machine learning fundamentals, model parallelism, and inference serving techniques.
  • Proficiency in Python; C++ experience is optional, for simulator design and data analysis.
  • At least 3 years of hands-on experience evaluating AI/ML workloads or performing performance analysis, modeling, and optimization for AI.
  • Experience defining metrics, designing experiments, and visualizing large performance datasets to identify resource bottlenecks.
  • Experience with deep learning frameworks such as PyTorch, TRT-LLM, vLLM, and SGLang.
  • A growth mindset and a pragmatic measure, iterate, and deliver approach.

Preferred Qualifications

  • A proven track record of working in multifunctional teams spanning algorithms, software, and hardware architecture.
  • Ability to distill complex analyses into clear recommendations for technical and non-technical collaborators.
  • Experience with GPU computing and CUDA.

Compensation and Benefits

  • Base salary for Level 3: USD 152,000–241,500 per year.
  • Base salary for Level 4: USD 184,000–287,500 per year.
  • Eligible employees may also receive equity and benefits.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs