Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Computer Vision
Distributed Systems @ 7
GPU @ 4
Helm @ 4
Kubernetes @ 4
LLM @ 7
Leadership @ 6
Machine Learning
Microservices
Observability @ 4
Prometheus
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's open-source benchmarking platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and enterprises use AIPerf to inform decisions on production inference, including choosing GPUs, optimizing costs, reducing latency, improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA's Dynamo organization and advance AIPerf for datacenter, local, and edge use cases spanning LLM, multimodal, diffusion, and computer vision inference.
This position combines hands-on leadership with expertise in systems engineering, inference infrastructure, and open-source communities. It has a direct effect on how AI performance is measured and improved.
Responsibilities
- Drive the technical roadmap for AIPerf's core infrastructure, including load generation, ZMQ-based microservices, GPU telemetry using DCGM and PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.
- Own the accuracy and statistical soundness of benchmark results used by engineering groups throughout the industry to inform production infrastructure decisions.
- Advise upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.
- Hire, mentor, and grow a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of overall software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
- 3+ years of engineering leadership experience as a tech lead, Technical Lead Manager, or engineering manager.
- Deep understanding of LLM inference mechanics, including TTFT, ITL, KV caching, prefill/decode, and speculative decoding, with the ability to reason about measurement correctness and reproducibility.
- Proven track record of collaborating across multifunctional groups and delivering production-quality output in high-velocity, high-external-visibility environments.
Preferred Qualifications
- Extensive experience with vLLM, TRT-LLM, or SGLang internals, along with contributions to their upstream projects.
- Experience building Kubernetes-native infrastructure, including operators, Helm charts, and GPU observability tooling such as DCGM, dcgm-exporter, and PyNVML.
- Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.
- History of leading or making meaningful contributions to active open-source projects with external communities.
Benefits
NVIDIA offers highly competitive salaries, a comprehensive benefits package, equity, and benefits for employees and their families. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications for this job will be accepted at least until June 1, 2026.