Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Algorithms
CUDA @ 5
Deep Learning @ 3
GPU
GenAI
Generative AI @ 3
JAX @ 3
LLM @ 3
Performance Analysis @ 3
PyTorch @ 3
Python @ 6
Robotics @ 3
SGLang @ 3
Software Development @ 3
TensorFlow @ 3
TensorRT @ 3
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Deep Learning Software Engineer passionate about analyzing and improving the performance of NVIDIA’s inference ecosystem. The team develops GPU-accelerated deep learning inference software, including TensorRT, deep learning benchmarking software, and performant solutions for deploying and serving models.
The role involves integrating TensorRT into open-source frameworks such as TensorRT-EdgeLLM and PyTorch, identifying performance opportunities, and optimizing state-of-the-art models across NVIDIA accelerators ranging from data center GPUs to edge systems. Responsibilities also include implementing graph compiler algorithms, frontend operators, and code generators across NVIDIA’s inference ecosystem, as well as collaborating on workflow improvements, performance modeling, performance analysis, kernel development, and inference software development.
Responsibilities
- Establish performance benchmarking methodologies and analysis workflows for NVIDIA’s inference ecosystem, including TensorRT, TensorRT-EdgeLLM, and Torch-TensorRT.
- Identify performance issues and optimization opportunities.
- Contribute features and code to NVIDIA and open-source inference frameworks, including TensorRT, TensorRT-EdgeLLM, and Torch-TensorRT.
- Develop optimized model pipelines involving quantization, scheduling, memory management, and distributed inference.
- Collaborate with teams across generative AI, automotive, robotics, image understanding, and speech understanding.
- Scale deep learning model performance across different NVIDIA architectures and accelerator types.
Requirements
- Bachelor’s, master’s, PhD, or equivalent experience in a relevant field such as Computer Science, Computer Engineering, EECS, or AI.
- Two years of relevant software development experience.
- Strong C++ and Python programming and software engineering skills.
- Experience with deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX.
- Experience with inference libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, or FlashInfer.
- Experience with performance analysis and optimization.
Preferred Qualifications
- Strong foundation and architectural knowledge of GPUs.
- Deep understanding of modern deep learning models and workloads, including Transformers, recommenders, automatic speech recognition, text-to-speech, and visual understanding.
- Proficiency in a deep learning programming domain-specific language such as CUDA, TileIR, CuTeDSL, CUTLASS, or Triton.
- Contributions to major LLM inference frameworks such as vLLM.
- Experience with deep learning inference graph compilers such as TorchDynamo or TorchInductor.
- Experience optimizing low-latency, resource-constrained systems or embedded AI pipelines, including Jetson systems or other edge AI accelerators.
Compensation and Benefits
The base salary range is USD 124,000–195,500 for Level 2 and USD 152,000–241,500 for Level 3. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 24, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.