Senior Deep Learning Software Engineer, Inference and Model Optimization
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms @ 7
CUDA @ 4
Communication @ 6
Debugging @ 6
Deep Learning @ 6
GPU @ 4
GenAI
Generative AI
JAX @ 6
LLM
Machine Learning @ 7
Mathematics @ 4
Performance Analysis @ 6
PyTorch @ 7
Python @ 7
TensorRT @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Algorithmic Model Optimization Team focuses on optimizing generative AI models, including large language models (LLMs) and diffusion models, for maximum inference efficiency. The team uses techniques such as neural architecture search, pruning, sparsity, quantization, and automated deployment strategies. Its work combines applied research with development of the TRT Model Optimizer software platform, which is used internally at NVIDIA and by external research and engineering teams.
The Senior Deep Learning Software Engineer will develop and scale automated inference and deployment solutions, working across machine learning frameworks, software architecture, and high-performance GPU kernel implementations.
Responsibilities
- Train, develop, and deploy generative AI models such as LLMs and diffusion models using NVIDIA's AI software stack.
- Use and extend the PyTorch 2.0 ecosystem, including TorchDynamo,
torch.export, andtorch.compile, to analyze and extract standardized model graph representations from arbitrary PyTorch models. - Develop high-performance inference optimization techniques, including automated model sharding, tensor parallelism, sequence parallelism, and efficient attention kernels with KV caching.
- Collaborate with teams across NVIDIA to integrate performant kernel implementations into the automated deployment solution.
- Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities.
- Improve inference performance so that NVIDIA's inference software solutions, including TensorRT, TRT-LLM, and TRT Model Optimizer, maintain and increase their market leadership.
- Architect and design a modular, scalable software platform offering broad model support and optimization techniques.
Requirements
- Master's degree, PhD, or equivalent experience in Computer Science, AI, Applied Mathematics, or a related field.
- At least 5 years of relevant work or research experience in deep learning.
- Excellent software design skills, including debugging, performance analysis, and test design.
- Strong proficiency in Python, PyTorch, and related machine learning tools such as Hugging Face.
- Strong algorithms and programming fundamentals.
- Good written and verbal communication skills, with the ability to work independently and collaboratively in a fast-paced environment.
Preferred Qualifications
- Contributions to PyTorch, JAX, or other machine learning frameworks.
- Knowledge of GPU architecture and the compilation stack, including the ability to understand and debug end-to-end performance.
- Familiarity with NVIDIA deep learning SDKs such as TensorRT.
- Experience writing high-performance GPU kernels for machine learning workloads using technologies such as CUDA, CUTLASS, or Triton.
Benefits
- Competitive base salary.
- Equity and benefits package.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 2, 2026.