Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 7
Deep Learning @ 7
GPU @ 7
JAX @ 7
LLM @ 6
Machine Learning @ 7
Performance Optimization @ 7
PyTorch @ 7
Python @ 7
SGLang @ 4
TensorFlow @ 7
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We're looking for outstanding AI systems engineers to develop technologies in the inference systems software stack. The team builds AI systems software to accelerate AI inference, including libraries, code generators, and GPU kernel technologies for NVIDIA hardware architectures. This includes new abstractions, efficient attention kernel implementations, LLM inference runtime components, and kernel code generators for large language models, agents, and other AI workloads.
Responsibilities
- Innovate and develop new AI systems technologies for efficient inference.
- Design, implement, and optimize kernels for high-impact AI workloads.
- Design and implement extensible abstractions for LLM serving engines.
- Build efficient just-in-time domain-specific compilers and runtimes.
- Collaborate with engineers across deep learning frameworks, libraries, kernels, and GPU architecture teams.
- Contribute to open-source communities such as FlashInfer, vLLM, and SGLang.
Requirements
- Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience. A PhD is preferred.
- 6+ years of academic or industry experience with machine learning or deep learning systems development is preferred.
- Strong experience developing or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX. Experience with inference engines and runtimes such as vLLM, SGLang, and MLC is ideal.
- Strong Python and C/C++ programming skills.
Preferred Qualifications
- Background in domain-specific compiler and library solutions for LLM inference and training, such as FlashInfer and Flash Attention.
- Expertise in inference engines such as vLLM and SGLang.
- Expertise in machine learning compilers such as Apache TVM and MLIR.
- Strong experience in GPU kernel development and performance optimization, especially with CUDA C/C++, cuTile, Triton, or similar technologies.
- Open-source project ownership or contributions.
Compensation and Benefits
The base salary range is USD 184,000–287,500, determined by location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until June 6, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.