Engineering Manager, Inference Benchmarking — AI Perf

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

Distributed Systems @ 4 GPU @ 4 Helm @ 4 Kubernetes @ 4 LLM @ 7 Leadership @ 6 Machine Learning Mentoring Microservices Observability @ 4 Prometheus SGLang @ 4 vLLM @ 4

Details

What you’ll be doing

  • Driving the technical roadmap for AIPerf's core infrastructure: load generation, ZMQ-based microservices, GPU telemetry (DCGM/PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.
  • Taking ownership for the accuracy and statistical soundness of benchmark results that engineering groups throughout the industry depend on to inform production infrastructure decisions.
  • Advising upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.
  • Hiring, mentoring, and growing a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 8+ overall years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
  • 3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager.
  • Deep understanding of LLM inference mechanics — TTFT, ITL, KV caching, Prefill/Decode, speculative decoding — and the ability to reason about measurement correctness and reproducibility.
  • Proven track record of collaborating across multi-functional groups and delivering production-quality output in high-velocity, high-external-visibility environments.

Ways to stand out from the crowd

  • Extensive experience with vLLM, TRT-LLM or SGLang internals along with contributions to their upstream projects.
  • Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling (DCGM, dcgm-exporter, PyNVML).
  • Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.
  • History leading or making meaningful contributions to active open-source projects with external communities.

More jobs at Nvidia

Similar jobs