Senior Deep Learning Frameworks CUDA Software Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 6 Communication @ 4 Deep Learning @ 4 GPU HPC @ 4 JAX @ 4 LLM @ 4 Performance Analysis PyTorch @ 4 Python @ 6 SGLang @ 4 System Architecture @ 4 vLLM @ 4

Details

Responsibilities

  • Integrate new CUDA features and Runtime abstractions in AI frameworks: from PoC to performance analysis to production
  • Perform deep analysis of AI workloads and frameworks to identify requirements and opportunities to innovate in the lower layers of the stack. Collaborate hands-on with teams working on the latest AI models
  • Own and drive improvements in the AI Compiler-Runtime interface to build speed-of-light multi-GPU multi-node solutions
  • Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads
  • Influence the roadmap of core CUDA to facilitate building next-gen DL frameworks
  • Collaborate with a very dynamic team across multiple time zones
  • Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and frameworks that enhance performance and programmability
  • Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning
  • Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products

Requirements

  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience)
  • 8+ years of relevant industry experience or equivalent academic experience after completed degree
  • Development experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang
  • Rapid prototyping and development with Python, C++, CUDA or related DSLs
  • Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)
  • Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)
  • Understanding of HPC/AI communication concepts
  • Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
  • Adaptability and passion to learn new frameworks and tools
  • Flexibility to work and communicate effectively across different teams and timezones

Benefits

  • Eligible for equity and benefits

More jobs at Nvidia

Similar jobs