Senior Deep Learning Software Engineer, Inference and Model Optimization

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Algorithms @ 7 CUDA @ 4 Communication @ 6 Debugging @ 6 Deep Learning @ 6 GPU @ 4 GenAI Generative AI JAX @ 6 LLM Machine Learning @ 7 Mathematics @ 4 Performance Analysis @ 6 PyTorch @ 7 Python @ 7 TensorRT @ 3

Details

NVIDIA's Algorithmic Model Optimization Team focuses on optimizing generative AI models, including large language models (LLMs) and diffusion models, for maximum inference efficiency. The team uses techniques such as neural architecture search, pruning, sparsity, quantization, and automated deployment strategies. Its work combines applied research with development of the TRT Model Optimizer software platform, which is used internally at NVIDIA and by external research and engineering teams.

The Senior Deep Learning Software Engineer will develop and scale automated inference and deployment solutions, working across machine learning frameworks, software architecture, and high-performance GPU kernel implementations.

Responsibilities

  • Train, develop, and deploy generative AI models such as LLMs and diffusion models using NVIDIA's AI software stack.
  • Use and extend the PyTorch 2.0 ecosystem, including TorchDynamo, torch.export, and torch.compile, to analyze and extract standardized model graph representations from arbitrary PyTorch models.
  • Develop high-performance inference optimization techniques, including automated model sharding, tensor parallelism, sequence parallelism, and efficient attention kernels with KV caching.
  • Collaborate with teams across NVIDIA to integrate performant kernel implementations into the automated deployment solution.
  • Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities.
  • Improve inference performance so that NVIDIA's inference software solutions, including TensorRT, TRT-LLM, and TRT Model Optimizer, maintain and increase their market leadership.
  • Architect and design a modular, scalable software platform offering broad model support and optimization techniques.

Requirements

  • Master's degree, PhD, or equivalent experience in Computer Science, AI, Applied Mathematics, or a related field.
  • At least 5 years of relevant work or research experience in deep learning.
  • Excellent software design skills, including debugging, performance analysis, and test design.
  • Strong proficiency in Python, PyTorch, and related machine learning tools such as Hugging Face.
  • Strong algorithms and programming fundamentals.
  • Good written and verbal communication skills, with the ability to work independently and collaboratively in a fast-paced environment.

Preferred Qualifications

  • Contributions to PyTorch, JAX, or other machine learning frameworks.
  • Knowledge of GPU architecture and the compilation stack, including the ability to understand and debug end-to-end performance.
  • Familiarity with NVIDIA deep learning SDKs such as TensorRT.
  • Experience writing high-performance GPU kernels for machine learning workloads using technologies such as CUDA, CUTLASS, or Triton.

Benefits

  • Competitive base salary.
  • Equity and benefits package.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 2, 2026.

More jobs at Nvidia

Similar jobs