Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Deep Learning @ 4
Distributed Systems @ 4
Experimentation
GPU @ 4
HPC @ 4
InfiniBand @ 7
Kubernetes @ 7
Leadership @ 6
Machine Learning
NCCL @ 4
Networking @ 7
Observability @ 6
Performance Monitoring
Profiling @ 4
PyTorch @ 4
Python @ 6
Slurm @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Deep Learning Software Infrastructure Engineer to support its Autonomous Vehicles project. The role focuses on building and scaling training libraries and infrastructure for end-to-end autonomous driving models, enabling training across thousands of GPUs and massive datasets while improving iteration speed and safety. The engineer will work closely with research and platform teams across NVIDIA.
Responsibilities
- Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand-GPU clusters.
- Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring.
- Build robust training pipelines and libraries for massive video datasets and rapid experimentation.
- Collaborate with researchers, model engineers, and internal platform teams to improve efficiency, minimize stalls, and increase training availability.
- Own core infrastructure components, including orchestration libraries, distributed training frameworks, and fault-resilient training systems.
- Partner with leadership to ensure infrastructure scales with growing GPU capacity and dataset size while maintaining developer efficiency and stability.
Requirements
- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, a related field, or equivalent experience.
- 12 or more years of professional experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure.
- Extensive knowledge of deep learning frameworks, preferably PyTorch; large-scale training technologies including DDP, FSDP, NCCL, tensor parallelism, and pipeline parallelism; and performance profiling.
- Strong systems background, including datacenter networking such as RoCE and InfiniBand, parallel filesystems such as Lustre, storage systems, and schedulers such as Slurm and Kubernetes.
- Proficiency in Python, including experience writing production-grade libraries, orchestration layers, and automation tools.
- Ability to work with multifunctional teams including ML researchers, infrastructure engineers, and product leads, and translate requirements into robust systems.
Preferred Qualifications
- Experience scaling large GPU training clusters with more than 1,000 GPUs.
- Expertise in fault resilience and high availability, including elastic training and large-scale observability.
- Demonstrated leadership as a hands-on technical authority, including encouraging others and establishing guidelines for ML systems engineering.
Compensation and Benefits
- Base salary range of USD 224,000–356,500 for Level 5.
- Base salary range of USD 272,000–431,250 for Level 6.
- Eligibility for equity and benefits.
- Applications accepted at least until August 10, 2026.
- NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.
- NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior AI Infrastructure Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year