Software Engineer, CUDA Deep Learning Systems

at Nvidia
USD 124,000-195,500 per year
MIDDLE SENIOR
✅ On-site

Tech Stack

AI @ 3 Agentic AI @ 3 Algorithms CUDA @ 3 Communication @ 3 Deep Learning @ 3 GPU GenAI Generative AI @ 3 JAX @ 3 LLM MPI @ 3 Machine Learning @ 3 NCCL @ 3 Performance Optimization @ 3 Profiling @ 2 PyTorch @ 3 Python @ 6 SGLang @ 3 TensorRT @ 3 vLLM @ 3

Details

We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and deep learning systems. The team explores and prototypes next-generation ideas that bridge deep learning algorithms and CUDA, focusing on model optimization, custom kernel development, and cluster-scale AI systems design. The role involves maximizing hardware performance for emerging AI workloads, from a single GPU to supercomputer clusters.

Responsibilities

  • Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
  • Architect and optimize distributed computing systems that scale from a single node to massive, cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms that improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency, and programmability.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write clean, effective, and maintainable code so exploratory prototypes can transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products.

Requirements

  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • At least 2 years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python.
  • Solid background in deep learning fundamentals, with a focus on transformers.
  • Strong understanding of distributed computing principles, multi-node scaling, and the performance challenges of cluster-scale execution.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with deep learning accelerator architectures such as GPUs, along with hands-on CUDA programming, kernel optimization, and workload profiling experience.
  • Experience profiling and optimizing generative AI models, including large language models.
  • Research background in machine learning systems or adjacent fields, and experience profiling and optimizing vision models, generative AI architectures, or diffusion models.
  • Track record of initiative and willingness to investigate problems across the stack.

Preferred Qualifications

  • Deep expertise in the performance internals and execution graphs of major deep learning training and inference frameworks, such as PyTorch, JAX, TensorRT, vLLM, sgLang, NeMo, and Megatron.
  • Hands-on experience with communication libraries such as NCCL, MPI, and UCX.
  • Experience with distributed machine learning techniques, including pipeline, tensor, and expert parallelism.
  • Knowledge of numerical methods and low-precision arithmetic, including NVFP4, MXFP4, FP8, and INT8, and their impact on deep learning accuracy and performance.
  • Background in deep learning compilers and ML systems, including graph-level and code-generation tools such as Triton, XLA, and torch.compile.
  • Experience with highly parallel or reinforcement-learning-style simulation environments.
  • Experience designing and implementing agentic AI systems for complex systems and infrastructure problems.

Compensation and Benefits

  • Base salary range: $124,000–$195,500 USD per year, determined by location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications accepted at least until August 9, 2026.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs