AI Inference Performance Engineer - New College Grad 2026

at Nvidia
USD 124,000-241,500 per year
JUNIOR
✅ On-site

Tech Stack

AI Algorithms @ 3 CUDA @ 3 Debugging @ 3 Deep Learning @ 3 GPU @ 3 GenAI Generative AI HPC @ 3 JAX @ 3 Kubernetes @ 3 LLM @ 6 MPI @ 3 NCCL @ 3 Profiling @ 3 PyTorch @ 3 Python @ 6 SGLang @ 3 Software Development @ 3 TensorRT @ 3 vLLM @ 3

Details

The team optimizes and benchmarks generative AI inference on NVIDIA's latest accelerators, defining performance standards across language models, video generation, and speech workloads. The role works directly with TensorRT-LLM, SGLang, and vLLM to build tools that evaluate serving performance at scale, at the intersection of GPU performance engineering and public accountability.

Responsibilities

  • Drive industry benchmark results by owning the end-to-end optimization pipeline and implementing and integrating optimizations in quantization, scheduling, memory management, and distributed inference across TensorRT-LLM, SGLang, and vLLM.
  • Define and optimize next-generation inference workloads, including multi-turn coding, agentic workflows, and other emerging AI use cases.
  • Collaborate with framework and kernel teams to optimize large-scale LLM-MoE models, vision-language models, video diffusion models, recommendation systems, and speech workloads.
  • Design and optimize distributed inference execution from single GPUs to rack-scale clusters.
  • Apply roofline analysis and systematic profiling to identify bottlenecks across CUDA kernels, frameworks, and serving layers.
  • Contribute to TensorRT-LLM, vLLM, SGLang, and other open-source projects.
  • Partner with architecture, kernel, and compiler teams to shape GPU roadmaps using real workload data.
  • Raise the team's technical bar, drive cross-functional execution on tight benchmark timelines, and lead a world-class team.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • At least 2 years of relevant software development experience.
  • Strong Python or C++ programming, software design, and software engineering skills.
  • Expertise with a deep learning framework such as PyTorch or JAX.
  • Proven experience delivering measurable performance improvements in deep learning inference or high-performance systems.
  • Deep understanding of LLM and VLM architectures and inference mechanics, including attention, KV caching, batching strategies, decode-phase bottlenecks, speculative decoding, and disaggregated serving.

Preferred Qualifications

  • Experience with an LLM framework such as TensorRT-LLM, vLLM, or SGLang, or with a deep learning compiler in inference, deployment, algorithms, or implementation.
  • Experience with performance modeling, profiling, debugging, and code optimization for deep learning, HPC, or other high-performance applications.
  • Experience with scale-out inference orchestration using MPI, NCCL, or Kubernetes on large GPU clusters.
  • Expertise in kernel development with CUTLASS, cuteDSL, TileLang, or OpenAI Triton.
  • Experience with compiler or runtime paths such as torch.compile, graph lowering, or operator fusion.
  • Architectural knowledge of CPUs, GPUs, FPGAs, or other deep learning accelerators.
  • GPU programming experience with CUDA.
  • Experience leading ambiguous, high-impact technical programs across multiple teams under tight deadlines.

Benefits

The role includes eligibility for equity and benefits. Applications will be accepted at least until June 7, 2026. NVIDIA is an equal opportunity employer and is committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs