Senior Deep Learning Software Infrastructure Engineer

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ Remote

Tech Stack

AI Deep Learning @ 4 Distributed Systems @ 4 Experimentation GPU @ 4 HPC @ 4 InfiniBand @ 7 Kubernetes @ 7 Leadership @ 6 Machine Learning NCCL @ 4 Networking @ 7 Observability @ 6 Performance Monitoring Profiling @ 4 PyTorch @ 4 Python @ 6 Slurm @ 7

Details

NVIDIA is seeking a Deep Learning Software Infrastructure Engineer to support its Autonomous Vehicles project. The role focuses on building and scaling training libraries and infrastructure for end-to-end autonomous driving models, enabling training across thousands of GPUs and massive datasets while improving iteration speed and safety. The engineer will work closely with research and platform teams across NVIDIA.

Responsibilities

  • Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand-GPU clusters.
  • Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring.
  • Build robust training pipelines and libraries for massive video datasets and rapid experimentation.
  • Collaborate with researchers, model engineers, and internal platform teams to improve efficiency, minimize stalls, and increase training availability.
  • Own core infrastructure components, including orchestration libraries, distributed training frameworks, and fault-resilient training systems.
  • Partner with leadership to ensure infrastructure scales with growing GPU capacity and dataset size while maintaining developer efficiency and stability.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, a related field, or equivalent experience.
  • 12 or more years of professional experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure.
  • Extensive knowledge of deep learning frameworks, preferably PyTorch; large-scale training technologies including DDP, FSDP, NCCL, tensor parallelism, and pipeline parallelism; and performance profiling.
  • Strong systems background, including datacenter networking such as RoCE and InfiniBand, parallel filesystems such as Lustre, storage systems, and schedulers such as Slurm and Kubernetes.
  • Proficiency in Python, including experience writing production-grade libraries, orchestration layers, and automation tools.
  • Ability to work with multifunctional teams including ML researchers, infrastructure engineers, and product leads, and translate requirements into robust systems.

Preferred Qualifications

  • Experience scaling large GPU training clusters with more than 1,000 GPUs.
  • Expertise in fault resilience and high availability, including elastic training and large-scale observability.
  • Demonstrated leadership as a hands-on technical authority, including encouraging others and establishing guidelines for ML systems engineering.

Compensation and Benefits

  • Base salary range of USD 224,000–356,500 for Level 5.
  • Base salary range of USD 272,000–431,250 for Level 6.
  • Eligibility for equity and benefits.
  • Applications accepted at least until August 10, 2026.
  • NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.
  • NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs