Senior Deep Learning Systems Engineer, Datacenters

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI Bash CUDA @ 6 Computer Vision @ 4 Deep Learning @ 6 Docker @ 6 GPU @ 4 LLM Linux @ 6 NLP Networking @ 4 Performance Analysis @ 7 Performance Monitoring @ 4 Profiling @ 4 PyTorch @ 6 Python @ 4 Slurm @ 6 System Architecture @ 7 TensorFlow @ 6

Details

The Deep Learning Systems Engineer will analyze the performance and power consumption of deep learning applications on datacenter-class hardware and influence the design and optimization of datacenters. The role involves understanding how CPU, GPU, networking, and I/O relate to deep learning architectures for natural language processing, computer vision, autonomous driving, large language models, and other applications.

Responsibilities

  • Develop software infrastructure to characterize and analyze a broad range of deep learning applications.
  • Evolve cost-efficient datacenter architectures tailored to the needs of large language models (LLMs).
  • Develop analysis and profiling tools using Python, Bash, and C++ to measure key performance metrics of deep learning workloads running on NVIDIA systems.
  • Analyze system and software characteristics of deep learning applications.
  • Develop analysis tools and methodologies to measure key performance metrics and estimate potential efficiency improvements.

Requirements

  • Bachelor's degree in Electrical Engineering or Computer Science, or equivalent experience. A master's or PhD degree is preferred.
  • Eight or more years of relevant experience.
  • Experience in at least one of the following areas:
    • System software, including Linux operating systems, compilers, GPU kernels using CUDA, or deep learning frameworks such as PyTorch and TensorFlow.
    • Silicon architecture and performance modeling or analysis, including CPU, GPU, memory, or network architecture.
  • Programming experience in C/C++ and Python.
  • A deep understanding of computer system architecture and performance analysis, with demonstrated hands-on experience.
  • Demonstrated ability to work in virtual environments and independently own tasks from beginning to end.
  • Exposure to containerization platforms such as Docker and datacenter workload managers such as Slurm is a plus.

Preferred Qualifications

  • Background in system software, operating system intrinsics, GPU kernels using CUDA, or deep learning frameworks such as PyTorch and TensorFlow.
  • Experience with silicon performance monitoring or profiling tools such as perf, gprof, nvidia-smi, and DCGM.
  • In-depth performance modeling experience in CPU, GPU, memory, or network architecture.
  • Exposure to Docker and Slurm.
  • Experience working with multisite or multifunctional teams.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.

Applications will be accepted at least until May 11, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs