Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 7
Deep Learning @ 7
GPU @ 7
JAX @ 7
LLM @ 6
Machine Learning @ 6
Performance Optimization @ 7
PyTorch @ 7
Python @ 7
SGLang @ 7
TensorFlow @ 7
vLLM @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack. We build innovative AI systems software to accelerate AI inference. As a member of the team, you'll develop libraries, code generators, and GPU kernel technologies for NVIDIA's hardware architecture. This includes designing and building new abstractions, efficient attention kernel implementations, new LLM inference runtime components, and kernel code generators to accelerate large language models, agents, and other high-impact AI workloads.
Responsibilities
- Innovate and develop new AI systems technologies for efficient inference.
- Design, implement, and optimize kernels for high-impact AI workloads.
- Design and implement extensible abstractions for LLM serving engines.
- Build efficient just-in-time domain-specific compilers and runtimes.
- Collaborate closely with engineers across deep learning frameworks, libraries, kernels, and GPU architecture teams.
- Contribute to open-source communities such as FlashInfer, vLLM, and SGLang.
Requirements
- Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience; PhD preferred.
- At least 6 years of academic or industry experience with ML/DL systems development preferred.
- Strong experience developing or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX, ideally including inference engines and runtimes such as vLLM, SGLang, and MLC.
- Strong Python and C/C++ programming skills.
- Strong experience in GPU kernel development and performance optimization, especially using CUDA C/C++, cuTile, Triton, or similar technologies, with hands-on experience in matrix multiplication.
Preferred Qualifications
- Background in domain-specific compiler and library solutions for LLM inference and training, such as FlashInfer and Flash Attention.
- Expertise in inference engines such as vLLM and SGLang.
- Expertise in machine learning compilers such as Apache TVM and MLIR.
- Open-source project ownership or contributions.
Compensation and Benefits
The base salary range is $184,000–$287,500 USD, determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until July 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.