Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 6
CUDA @ 3
Communication @ 6
Data Structures @ 6
Debugging @ 6
Deep Learning @ 6
Distributed Systems @ 6
GPU @ 3
GitHub @ 6
Go @ 7
LLM @ 4
LLVM @ 4
Machine Learning
NCCL @ 3
Parallel Programming @ 6
Performance Optimization
Profiling @ 6
PyTorch @ 4
Python @ 7
Reinforcement Learning @ 4
Rust @ 7
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking highly skilled and motivated software engineers to build AI inference systems that serve large-scale models with extreme efficiency. You will architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You will collaborate across inference performance, kernels, training, large-scale serving, and research teams to advance accelerated computing for AI.
Responsibilities
- Contribute features to vLLM that support the newest models, NVIDIA GPU hardware features, and serving runtime algorithms.
- Profile and optimize the vLLM inference framework using techniques such as speculative decoding, 5D parallelism, and prefill-decode disaggregation.
- Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.
- Conduct and publish original research that advances the Pareto frontier in ML systems.
- Survey recent publications and integrate research ideas and prototypes into production-grade, open-source software.
- Develop, optimize, and benchmark GPU kernels using hand-tuned and compiler-generated approaches, including fusion, autotuning, and memory/layout optimization.
Requirements
- Bachelor's, Master's, or PhD degree in Computer Science, Computer Engineering, or Software Engineering.
- At least 5 years of industry experience in software engineering or equivalent research experience.
- Strong programming skills in Python and one of C, C++, Go, or Rust.
- Solid computer science fundamentals, including algorithms and data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, and deep learning theory.
- Knowledge of and passion for performance engineering in ML frameworks such as PyTorch and inference engines such as vLLM and SGLang.
- Familiarity with GPU programming and performance, including CUDA, memory hierarchy, streams, and NCCL.
- Proficiency with profiling and debugging tools such as Nsight Systems and Nsight Compute.
- Excellent debugging, problem-solving, and communication skills, with the ability to work effectively in a fast-paced, multifunctional environment.
Preferred Qualifications
- Experience developing major features and optimizations for LLM inference engines such as vLLM and SGLang.
- Hands-on experience with LLM inference and training runtimes, including deploying LLMs to production, large-scale LLM pre-training, and reinforcement learning.
- Experience with ML compilers and DSLs such as Triton, CuTe, MLIR/LLVM, and XLA.
- Experience with GPU libraries such as CUTLASS and GPU features such as CUDA Graphs and Tensor Cores.
- Experience with speculative decoding training and runtime features, including tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, and EAGLE.
- Contributions to open-source projects and/or publications. Applicants are encouraged to include links to GitHub pull requests, published papers, and artifacts.
About the Team
The team consists of experts in AI, systems, and performance optimization. Its leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. NVIDIA's mission is to advance AI research and development and create technologies that enable people to harness the potential of AI.
Compensation and Benefits
- Base salary for Level 3: CAD 135,000–185,000 per year.
- Base salary for Level 4: CAD 170,000–220,000 per year.
- Eligibility for equity and benefits.
Applications will be accepted at least until August 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.