Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
CUDA @ 7
Engineering Management
GPU @ 7
LLM @ 4
Leadership @ 7
Python @ 7
SGLang @ 4
Technical Leadership @ 7
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is accelerating LLM inference across the stack and across open-source LLM frameworks such as TensorRT-LLM, vLLM, and SGLang. The Engineering Manager will lead the development of next-generation LLM, VLM, and VLA inference software technologies in a hands-on leadership role combining deep technical expertise with engineering management. The role involves collaboration with NVIDIA researchers, GPU architects, and teams across the company to deliver production-grade, high-performance software.
Responsibilities
- Lead and grow a team responsible for improving LLM inference performance across TensorRT-LLM, vLLM, SGLang, and Dynamo on NVIDIA data center products.
- Drive the design, implementation, and optimization of performance-critical LLM inference features.
- Improve LLM inference performance on current and upcoming NVIDIA data center architectures and GPUs.
- Improve inference performance for important foundation models.
- Work with inference benchmark teams to tune performance for key workloads.
- Integrate cutting-edge NVIDIA technologies and provide an intuitive developer experience for LLM deployment.
- Lead software development execution, including project planning, milestone delivery, and cross-functional coordination.
Requirements
- MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.
- 7+ years of overall software engineering experience, including 3+ years of technical leadership experience.
- Proven ability to lead and scale high-performing engineering teams, especially across distributed and cross-functional groups.
- Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries.
- Demonstrated expertise in large language models, vision language models, and/or inference in general.
Preferred Qualifications
- Deep understanding of GPU architecture, CUDA programming, and system-level performance tuning.
- Background in LLM inference or experience with frameworks such as TensorRT-LLM, vLLM, or SGLang.
- Passion for building scalable, user-friendly APIs and enabling developers in the AI ecosystem.
- Proven track record of growing and managing teams that encourage idea sharing, empower team members, and provide opportunities for professional growth.
Compensation and Benefits
The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 224,000–356,500 for Level 3 and USD 272,000–431,250 for Level 4. The role is also eligible for equity and benefits.
Applications will be accepted at least until August 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.