Engineering Manager, LLM Performance

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 4 API @ 6 CUDA @ 7 Engineering Management GPU @ 7 LLM @ 4 Leadership @ 7 Python @ 7 SGLang @ 4 Technical Leadership @ 7 TensorRT @ 4 vLLM @ 4

Details

NVIDIA is accelerating LLM inference across the stack and across open-source LLM frameworks such as TensorRT-LLM, vLLM, and SGLang. The Engineering Manager will lead the development of next-generation LLM, VLM, and VLA inference software technologies in a hands-on leadership role combining deep technical expertise with engineering management. The role involves collaboration with NVIDIA researchers, GPU architects, and teams across the company to deliver production-grade, high-performance software.

Responsibilities

  • Lead and grow a team responsible for improving LLM inference performance across TensorRT-LLM, vLLM, SGLang, and Dynamo on NVIDIA data center products.
  • Drive the design, implementation, and optimization of performance-critical LLM inference features.
  • Improve LLM inference performance on current and upcoming NVIDIA data center architectures and GPUs.
  • Improve inference performance for important foundation models.
  • Work with inference benchmark teams to tune performance for key workloads.
  • Integrate cutting-edge NVIDIA technologies and provide an intuitive developer experience for LLM deployment.
  • Lead software development execution, including project planning, milestone delivery, and cross-functional coordination.

Requirements

  • MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.
  • 7+ years of overall software engineering experience, including 3+ years of technical leadership experience.
  • Proven ability to lead and scale high-performing engineering teams, especially across distributed and cross-functional groups.
  • Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries.
  • Demonstrated expertise in large language models, vision language models, and/or inference in general.

Preferred Qualifications

  • Deep understanding of GPU architecture, CUDA programming, and system-level performance tuning.
  • Background in LLM inference or experience with frameworks such as TensorRT-LLM, vLLM, or SGLang.
  • Passion for building scalable, user-friendly APIs and enabling developers in the AI ecosystem.
  • Proven track record of growing and managing teams that encourage idea sharing, empower team members, and provide opportunities for professional growth.

Compensation and Benefits

The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 224,000–356,500 for Level 3 and USD 272,000–431,250 for Level 4. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs