Senior Deep Learning Engineer – Autonomous Vehicles

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Deep Learning @ 4 Distributed Systems @ 8 Experimentation GPU @ 4 HPC InfiniBand @ 7 Kubernetes @ 7 Leadership @ 6 Machine Learning @ 8 NCCL @ 4 Networking @ 7 Observability @ 6 Performance Monitoring Profiling @ 4 PyTorch @ 4 Python @ 6 Reinforcement Learning @ 3 Slurm @ 7 Technical Leadership @ 6

Details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI to define the next era of computing, with GPUs powering computers, robots, and self-driving cars that can understand the world.

The Autonomous Vehicles team is seeking a Senior Deep Learning Systems Engineer to build and scale training libraries and infrastructure for end-to-end autonomous driving models. This role will enable training across thousands of GPUs and massive datasets, accelerating iteration speed and improving safety while working closely with research and platform teams across NVIDIA.

Responsibilities

  • Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand-GPU clusters.
  • Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring.
  • Build robust training pipelines and libraries for massive video datasets and rapid experimentation.
  • Collaborate with researchers, model engineers, and internal platform teams to improve efficiency, minimize stalls, and increase training availability.
  • Own core infrastructure components, including orchestration libraries, distributed training frameworks, and fault-resilient training systems.
  • Partner with leadership to ensure infrastructure scales with increasing GPU capacity and dataset size while maintaining developer efficiency and stability.

Requirements

  • Bachelor's, master's, or doctoral degree in Computer Science, Electrical or Computer Engineering, or a related field, or equivalent experience.
  • At least 12 years of professional experience building and scaling high-performance distributed systems, ideally in machine learning, high-performance computing, or large-scale data infrastructure.
  • Extensive knowledge of deep learning frameworks, preferably PyTorch, large-scale training technologies such as DDP, FSDP, NCCL, and tensor or pipeline parallelism, and performance profiling.
  • Strong systems background, including datacenter networking such as RoCE and InfiniBand, parallel filesystems such as Lustre, storage systems, and schedulers such as Slurm and Kubernetes.
  • Proficiency in Python and C++, with experience writing production-grade libraries, orchestration layers, and automation tools.
  • Ability to work closely with multifunctional teams, including machine learning researchers, infrastructure engineers, and product leads, and translate requirements into robust systems.

Preferred Qualifications

  • Experience scaling large GPU training clusters with more than 1,000 GPUs.
  • Contributions to open-source machine learning systems libraries, such as PyTorch, NCCL, FSDP, schedulers, or storage clients.
  • Expertise in fault resilience and high availability, including elastic training and large-scale observability.
  • Hands-on technical leadership skills, including establishing guidelines for machine learning systems engineering.
  • Familiarity with reinforcement learning at scale, particularly for simulation-heavy workloads.

Compensation and Benefits

The base salary range is $224,000–$356,500 USD, determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until September 12, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs