Senior Software Engineer, CUDA Deep Learning Systems

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic AI @ 4 Algorithms CUDA @ 4 Communication @ 4 Deep Learning @ 4 GPU @ 3 GenAI Generative AI @ 4 JAX @ 6 LLM MPI @ 4 Machine Learning @ 4 NCCL @ 4 Performance Optimization @ 4 Profiling @ 4 PyTorch @ 6 Python @ 7 SGLang @ 6 TensorRT @ 6 vLLM @ 6

Details

We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and deep learning systems. The team explores and prototypes next-generation ideas that bridge deep learning algorithms and CUDA, with a focus on modern accelerator architectures, model optimization, custom kernel development, and cluster-scale AI systems.

This research-oriented role focuses on maximizing hardware performance for emerging AI workloads, from single GPUs to supercomputer clusters. Exploratory prototypes may transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products.

Responsibilities

  • Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
  • Architect and optimize distributed computing systems that scale from a single node to massive, cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
  • Improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency, and programmability.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write clean, effective, and maintainable code.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python.
  • Solid understanding of deep learning fundamentals, particularly transformers.
  • Strong understanding of distributed computing, multi-node scaling, and the performance challenges of cluster-scale execution.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with GPU and other deep learning accelerator architectures.
  • Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
  • Experience profiling and optimizing generative AI models, including large language models.
  • Research background in machine learning systems or adjacent fields, with experience profiling and optimizing vision models, generative AI architectures, or diffusion models.
  • Track record of initiative and willingness to investigate problems across the stack.

Preferred Qualifications

  • Deep expertise in the performance internals and execution graphs of deep learning training and inference frameworks such as PyTorch, JAX, TensorRT, vLLM, sgLang, NeMo, or Megatron.
  • Hands-on experience with communication libraries such as NCCL, MPI, or UCX.
  • Experience with distributed machine learning techniques, including pipeline, tensor, and expert parallelism.
  • Knowledge of numerical methods and low-precision arithmetic, including NVFP4, MXFP4, FP8, and INT8, and their impact on deep learning accuracy and performance.
  • Background in deep learning compilers and ML systems, including graph-level and code-generation tools such as Triton, XLA, and torch.compile.
  • Experience with highly parallel or reinforcement-learning-style simulation environments.
  • Experience designing and implementing agentic AI systems for complex systems and infrastructure problems.

Compensation and Benefits

  • Base salary for Level 4: USD 184,000–287,500 per year.
  • Base salary for Level 5: USD 224,000–356,500 per year.
  • Eligibility for equity and benefits.
  • Applications will be accepted at least until August 9, 2026.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs