Inference Performance Engineer, AI Inference Configuration Optimization

at Nvidia
USD 124,000-241,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 4 Agentic AI @ 4 CUDA @ 7 Communication @ 7 GPU @ 4 LLM @ 6 Mathematics @ 4 NCCL @ 4 Profiling @ 4 PyTorch @ 4 Python @ 7 SGLang @ 6 TensorRT @ 6 vLLM @ 6

Details

NVIDIA is seeking an Inference Performance Engineer to optimize large-scale AI inference benchmarks through an autonomous optimization framework. The role focuses on extracting maximum performance from GPUs by designing techniques that AI agents can repeatedly benchmark, profile, and tune.

Responsibilities

  • Distill performance expertise into reusable skills, workflows, and evidence-backed methodologies for autonomous AI agents.
  • Review agent-generated experiments, validate findings, and curate best-known configurations.
  • Improve AI inference workloads by optimizing throughput per GPU and user interactivity through configuration options, parallelism, batching, KV cache handling, quantization, and speculative decoding.
  • Measure and optimize aggregated and disaggregated serving architectures using TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA GPU platforms.
  • Profile workloads with Nsight Systems, kernel traces, and internal analysis tools.
  • Apply roofline and speed-of-light analysis to identify performance headroom and drive improvements from hypothesis through measured results.
  • Contribute upstream serving framework patches, optimized kernels, and deployment recipes while maintaining strict model correctness.
  • Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to deliver performance improvements.

Requirements

  • Bachelor's, master's, or doctoral degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field, or equivalent experience.
  • 3 or more years of relevant engineering experience.
  • Extensive knowledge of AI model execution efficiency and optimization, including continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing, mixture-of-experts serving, quantization, and serving service-level agreements.
  • Hands-on experience benchmarking and profiling GPU workloads with tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch Profiler.
  • Ability to interpret kernel-level performance data.
  • Strong Python engineering skills and ability to navigate and modify large C++/CUDA serving codebases.
  • Rigorous experimental methodology using controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.
  • Strong written and verbal communication skills.

Preferred Qualifications

  • Contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.
  • Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.
  • CUDA kernel authorship or optimization experience on Hopper or Blackwell architectures, including Tensor Cores, TMA, and warp specialization.
  • Results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.
  • Experience building or operating agentic AI workflows to automate engineering tasks.

Benefits

NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. The position is full time. Applications will be accepted at least until August 10, 2026.

More jobs at Nvidia

Similar jobs