Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Distributed Systems @ 4
GPU @ 4
Helm @ 4
Kubernetes @ 4
LLM @ 7
Leadership @ 6
Machine Learning
Mentoring
Microservices
Observability @ 4
Prometheus
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
What you’ll be doing
- Driving the technical roadmap for AIPerf's core infrastructure: load generation, ZMQ-based microservices, GPU telemetry (DCGM/PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.
- Taking ownership for the accuracy and statistical soundness of benchmark results that engineering groups throughout the industry depend on to inform production infrastructure decisions.
- Advising upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.
- Hiring, mentoring, and growing a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
- 8+ overall years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
- 3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager.
- Deep understanding of LLM inference mechanics — TTFT, ITL, KV caching, Prefill/Decode, speculative decoding — and the ability to reason about measurement correctness and reproducibility.
- Proven track record of collaborating across multi-functional groups and delivering production-quality output in high-velocity, high-external-visibility environments.
Ways to stand out from the crowd
- Extensive experience with vLLM, TRT-LLM or SGLang internals along with contributions to their upstream projects.
- Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling (DCGM, dcgm-exporter, PyNVML).
- Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.
- History leading or making meaningful contributions to active open-source projects with external communities.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Perplexity AI · New York City, United States, Ireland, London, United Kingdom, San Francisco, United States
USD 250,000-485,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer - Vlm Microservices for Neural Reconstruction
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Engineer, AI Tooling and Workflows
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Security Engineer, AI Security
Reddit · United States
USD 190,800-267,100 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year