Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Deep Learning @ 4
Distributed Systems @ 8
Experimentation
GPU @ 4
HPC
InfiniBand @ 7
Kubernetes @ 7
Leadership @ 6
Machine Learning @ 8
NCCL @ 4
Networking @ 7
Observability @ 6
Performance Monitoring
Profiling @ 4
PyTorch @ 4
Python @ 6
Reinforcement Learning @ 3
Slurm @ 7
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI to define the next era of computing, with GPUs powering computers, robots, and self-driving cars that can understand the world.
The Autonomous Vehicles team is seeking a Senior Deep Learning Systems Engineer to build and scale training libraries and infrastructure for end-to-end autonomous driving models. This role will enable training across thousands of GPUs and massive datasets, accelerating iteration speed and improving safety while working closely with research and platform teams across NVIDIA.
Responsibilities
- Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand-GPU clusters.
- Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring.
- Build robust training pipelines and libraries for massive video datasets and rapid experimentation.
- Collaborate with researchers, model engineers, and internal platform teams to improve efficiency, minimize stalls, and increase training availability.
- Own core infrastructure components, including orchestration libraries, distributed training frameworks, and fault-resilient training systems.
- Partner with leadership to ensure infrastructure scales with increasing GPU capacity and dataset size while maintaining developer efficiency and stability.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Electrical or Computer Engineering, or a related field, or equivalent experience.
- At least 12 years of professional experience building and scaling high-performance distributed systems, ideally in machine learning, high-performance computing, or large-scale data infrastructure.
- Extensive knowledge of deep learning frameworks, preferably PyTorch, large-scale training technologies such as DDP, FSDP, NCCL, and tensor or pipeline parallelism, and performance profiling.
- Strong systems background, including datacenter networking such as RoCE and InfiniBand, parallel filesystems such as Lustre, storage systems, and schedulers such as Slurm and Kubernetes.
- Proficiency in Python and C++, with experience writing production-grade libraries, orchestration layers, and automation tools.
- Ability to work closely with multifunctional teams, including machine learning researchers, infrastructure engineers, and product leads, and translate requirements into robust systems.
Preferred Qualifications
- Experience scaling large GPU training clusters with more than 1,000 GPUs.
- Contributions to open-source machine learning systems libraries, such as PyTorch, NCCL, FSDP, schedulers, or storage clients.
- Expertise in fault resilience and high availability, including elastic training and large-scale observability.
- Hands-on technical leadership skills, including establishing guidelines for machine learning systems engineering.
- Familiarity with reinforcement learning at scale, particularly for simulation-heavy workloads.
Compensation and Benefits
The base salary range is $224,000–$356,500 USD, determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 12, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.