Senior Inference Engineer, GPU Kernel Optimization

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic AI @ 4 Agentic Systems @ 4 CUDA @ 4 GPU @ 4 LLM @ 4 LLVM @ 7 Performance Analysis Profiling @ 4 Python @ 7 SGLang @ 4 vLLM @ 4

Details

NVIDIA is looking for a Senior Inference Engineer to join its LLM Inference Performance Analysis and Optimization team. The team develops silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. The role involves close collaboration with compiler, kernel, hardware, and framework organizations to identify bottlenecks and deliver measurable performance improvements for LLM inference.

Responsibilities

  • Build GPU kernel microbenchmarking systems that measure competing kernel implementations with real-silicon fidelity across the configuration space required by production LLM deployments.
  • Perform end-to-end model performance analysis by connecting performance data to model-level serving economics, identifying high-value optimization opportunities, and producing optimization policies for production inference deployments.
  • Develop agentic kernel optimization systems that use AI-driven analysis to diagnose performance gaps and explore optimization opportunities across the kernel ecosystem.
  • Validate optimization findings with rigorous silicon measurements.
  • Collaborate with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.

Requirements

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • Six or more years of relevant industry experience.
  • Experience building or directing agentic AI systems involving code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling experience with CUPTI, NSYS, and NCU, including the ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
  • Clear understanding of how kernel selection affects model-level throughput and latency.
  • Working knowledge of GPU kernel optimization using CUDA, CUTLASS, Triton, or equivalent technologies.
  • Ability to read PTX or SASS output.

Preferred Qualifications

  • Deep knowledge of SASS- and PTX-level kernel analysis, compiler middle-end optimization, or GPU code-generation pipelines such as LLVM, MLIR, and ptxas.
  • Experience shipping agentic systems end-to-end, including tool invention, multi-agent orchestration, and silicon-verified validation in a performance engineering or kernel optimization context.
  • Contributions to open-source LLM inference or GPU kernel libraries such as FlashInfer, Triton, or CUTLASS.

Benefits

  • Competitive base salary ranging from $184,000 to $287,500 per year, depending on location, experience, and compensation for similar positions.
  • Eligibility for equity and benefits.
  • NVIDIA provides a comprehensive benefits package.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications for this job will be accepted at least until July 31, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs