Senior DL Performance Efficiency Architect

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 6 GPU HPC LLM @ 7 Leadership @ 6 Performance Analysis @ 7 Performance Optimization @ 6 Technical Leadership @ 6

Details

NVIDIA's accelerated computing platform is enabling generational improvements in large language models, while the scale and complexity of these models create new challenges in computational efficiency. This role will drive a unified strategy for making LLMs more efficient from research through deployment by combining model innovation, systems expertise, and hardware awareness. The position will lead a multidisciplinary effort, establish technical direction for LLM efficiency, and help shape how future models and computing platforms are designed together.

Responsibilities

  • Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training, and inference systems.
  • Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, identifying opportunities for model-system-hardware co-design.
  • Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
  • Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps.

Requirements

  • Master's or PhD degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
  • Strong understanding of LLM architectures, training and inference workloads, and the tradeoffs between model quality, computational cost, memory footprint, latency, throughput, and power.
  • Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
  • Proven ability to provide technical leadership and drive complex optimization projects from concept to production.

Preferred Qualifications

  • Track record of delivering measurable improvements in throughput, cost per token, energy per token, memory efficiency, or time to train.
  • First-principles approach to improving LLM efficiency through measurement, modeling, optimization, and delivery.
  • Familiarity with low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding.
  • Experience co-designing model architectures with training, inference, compiler, or hardware constraints.
  • Experience influencing accelerator, system, or datacenter architecture based on future AI workload requirements.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs