Inference Performance Engineer, AI Inference Configuration Optimization
at Nvidia
USD 124,000-241,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agentic AI @ 4
CUDA @ 7
Communication @ 7
GPU @ 4
LLM @ 6
Mathematics @ 4
NCCL @ 4
Profiling @ 4
PyTorch @ 4
Python @ 7
SGLang @ 6
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an Inference Performance Engineer to optimize large-scale AI inference benchmarks through an autonomous optimization framework. The role focuses on extracting maximum performance from GPUs by designing techniques that AI agents can repeatedly benchmark, profile, and tune.
Responsibilities
- Distill performance expertise into reusable skills, workflows, and evidence-backed methodologies for autonomous AI agents.
- Review agent-generated experiments, validate findings, and curate best-known configurations.
- Improve AI inference workloads by optimizing throughput per GPU and user interactivity through configuration options, parallelism, batching, KV cache handling, quantization, and speculative decoding.
- Measure and optimize aggregated and disaggregated serving architectures using TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA GPU platforms.
- Profile workloads with Nsight Systems, kernel traces, and internal analysis tools.
- Apply roofline and speed-of-light analysis to identify performance headroom and drive improvements from hypothesis through measured results.
- Contribute upstream serving framework patches, optimized kernels, and deployment recipes while maintaining strict model correctness.
- Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to deliver performance improvements.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field, or equivalent experience.
- 3 or more years of relevant engineering experience.
- Extensive knowledge of AI model execution efficiency and optimization, including continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing, mixture-of-experts serving, quantization, and serving service-level agreements.
- Hands-on experience benchmarking and profiling GPU workloads with tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch Profiler.
- Ability to interpret kernel-level performance data.
- Strong Python engineering skills and ability to navigate and modify large C++/CUDA serving codebases.
- Rigorous experimental methodology using controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.
- Strong written and verbal communication skills.
Preferred Qualifications
- Contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.
- Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.
- CUDA kernel authorship or optimization experience on Hopper or Blackwell architectures, including Tensor Cores, TMA, and warp specialization.
- Results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.
- Experience building or operating agentic AI workflows to automate engineering tasks.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. The position is full time. Applications will be accepted at least until August 10, 2026.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Deep Learning Frameworks CUDA Software Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year