Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 3
CUDA @ 3
Debugging @ 3
Deep Learning @ 3
GPU @ 3
GenAI
Generative AI
HPC @ 3
JAX @ 3
Kubernetes @ 3
LLM @ 6
MPI @ 3
NCCL @ 3
Profiling @ 3
PyTorch @ 3
Python @ 6
SGLang @ 3
Software Development @ 3
TensorRT @ 3
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The team optimizes and benchmarks generative AI inference on NVIDIA's latest accelerators, defining performance standards across language models, video generation, and speech workloads. The role works directly with TensorRT-LLM, SGLang, and vLLM to build tools that evaluate serving performance at scale, at the intersection of GPU performance engineering and public accountability.
Responsibilities
- Drive industry benchmark results by owning the end-to-end optimization pipeline and implementing and integrating optimizations in quantization, scheduling, memory management, and distributed inference across TensorRT-LLM, SGLang, and vLLM.
- Define and optimize next-generation inference workloads, including multi-turn coding, agentic workflows, and other emerging AI use cases.
- Collaborate with framework and kernel teams to optimize large-scale LLM-MoE models, vision-language models, video diffusion models, recommendation systems, and speech workloads.
- Design and optimize distributed inference execution from single GPUs to rack-scale clusters.
- Apply roofline analysis and systematic profiling to identify bottlenecks across CUDA kernels, frameworks, and serving layers.
- Contribute to TensorRT-LLM, vLLM, SGLang, and other open-source projects.
- Partner with architecture, kernel, and compiler teams to shape GPU roadmaps using real workload data.
- Raise the team's technical bar, drive cross-functional execution on tight benchmark timelines, and lead a world-class team.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- At least 2 years of relevant software development experience.
- Strong Python or C++ programming, software design, and software engineering skills.
- Expertise with a deep learning framework such as PyTorch or JAX.
- Proven experience delivering measurable performance improvements in deep learning inference or high-performance systems.
- Deep understanding of LLM and VLM architectures and inference mechanics, including attention, KV caching, batching strategies, decode-phase bottlenecks, speculative decoding, and disaggregated serving.
Preferred Qualifications
- Experience with an LLM framework such as TensorRT-LLM, vLLM, or SGLang, or with a deep learning compiler in inference, deployment, algorithms, or implementation.
- Experience with performance modeling, profiling, debugging, and code optimization for deep learning, HPC, or other high-performance applications.
- Experience with scale-out inference orchestration using MPI, NCCL, or Kubernetes on large GPU clusters.
- Expertise in kernel development with CUTLASS, cuteDSL, TileLang, or OpenAI Triton.
- Experience with compiler or runtime paths such as torch.compile, graph lowering, or operator fusion.
- Architectural knowledge of CPUs, GPUs, FPGAs, or other deep learning accelerators.
- GPU programming experience with CUDA.
- Experience leading ambiguous, high-impact technical programs across multiple teams under tight deadlines.
Benefits
The role includes eligibility for equity and benefits. Applications will be accepted at least until June 7, 2026. NVIDIA is an equal opportunity employer and is committed to an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Deep Learning Software Engineer, Inference
Nvidia · United States
USD 152,000-287,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year