Engineering Manager, Inference Benchmarking — AI Perf

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI Computer Vision Distributed Systems @ 7 GPU @ 4 Helm @ 4 Kubernetes @ 4 LLM @ 7 Leadership @ 6 Machine Learning Microservices Observability @ 4 Prometheus SGLang @ 4 vLLM @ 4

Details

NVIDIA's open-source benchmarking platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and enterprises use AIPerf to inform decisions on production inference, including choosing GPUs, optimizing costs, reducing latency, improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA's Dynamo organization and advance AIPerf for datacenter, local, and edge use cases spanning LLM, multimodal, diffusion, and computer vision inference.

This position combines hands-on leadership with expertise in systems engineering, inference infrastructure, and open-source communities. It has a direct effect on how AI performance is measured and improved.

Responsibilities

  • Drive the technical roadmap for AIPerf's core infrastructure, including load generation, ZMQ-based microservices, GPU telemetry using DCGM and PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.
  • Own the accuracy and statistical soundness of benchmark results used by engineering groups throughout the industry to inform production infrastructure decisions.
  • Advise upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.
  • Hire, mentor, and grow a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of overall software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
  • 3+ years of engineering leadership experience as a tech lead, Technical Lead Manager, or engineering manager.
  • Deep understanding of LLM inference mechanics, including TTFT, ITL, KV caching, prefill/decode, and speculative decoding, with the ability to reason about measurement correctness and reproducibility.
  • Proven track record of collaborating across multifunctional groups and delivering production-quality output in high-velocity, high-external-visibility environments.

Preferred Qualifications

  • Extensive experience with vLLM, TRT-LLM, or SGLang internals, along with contributions to their upstream projects.
  • Experience building Kubernetes-native infrastructure, including operators, Helm charts, and GPU observability tooling such as DCGM, dcgm-exporter, and PyNVML.
  • Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.
  • History of leading or making meaningful contributions to active open-source projects with external communities.

Benefits

NVIDIA offers highly competitive salaries, a comprehensive benefits package, equity, and benefits for employees and their families. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications for this job will be accepted at least until June 1, 2026.

More jobs at Nvidia

Similar jobs