Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
CUDA @ 4
Computer Vision @ 6
Debugging
GPU @ 4
Leadership @ 7
Machine Learning @ 4
PyTorch @ 7
Python @ 7
Robotics @ 4
Technical Leadership @ 7
TensorRT @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a strong engineer to join the DRIVE Road Structure, Online Mapping, and Context Fusion team. In this role, you will help shape the future of NVIDIA's L3/L4 autonomous-driving solution by building a learned 3D/4D world model that fuses navigation, ego-motion, perception, and sensor signals. You will work closely with perception, prediction, planning, and simulation teams to deliver a complete, temporally consistent, uncertainty-aware world representation capable of supporting challenging roads and intersections.
The role focuses on developing a shared multimodal scene representation for road understanding, 3D object and occupancy perception, motion prediction, and route-conditioned planning.
Responsibilities
- Design and develop learning-based, multimodal sensor-fusion systems that transform synchronized sensor history, ego-motion, navigation context, and driving context into a unified spatiotemporal world representation.
- Build architectures that jointly reason over camera, LiDAR, radar, and vehicle-state inputs, including calibration, synchronization, coordinate transforms, sensor latency, and uncertainty handling.
- Develop end-to-end and multi-task models for road graph elements such as lanes, boundaries, crosswalks, and traffic controls; semantic scene understanding; and occupancy and free-space representations, including uncertain and occluded regions.
- Develop scalable multimodal fusion architectures, including Transformer-based early, late, and hierarchical fusion; BEV, point/voxel, and image-based representations; temporal context aggregation; and cross-modal attention.
- Create training, fine-tuning, and evaluation pipelines for large-scale multimodal datasets.
- Define multi-task objectives and metrics that balance perception quality, geometric consistency, prediction accuracy, latency, and safety-critical behavior.
- Investigate foundation-model approaches for autonomous driving, including vision-language models, multimodal pre-training, representation learning, and efficient deployment of learned world models.
- Work with perception, mapping, prediction, planning, simulation, data, and embedded-software teams to convert research advances into robust, production-quality autonomous-vehicle systems.
- Develop analysis and debugging tools for model failures, cross-sensor disagreement, long-tail scenarios, distribution shift, and regressions in closed-loop simulation and on-road evaluation.
Requirements
- BS, MS, or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a related technical field, or equivalent experience.
- 8+ years of experience, including at least 2+ years in the autonomous-vehicle or robotics industry and 2+ years of technical leadership experience.
- Strong experience developing production-quality sensor-fusion, perception, state-estimation, or autonomous-driving systems.
- Experience developing learning-based multimodal perception or fusion involving at least two of the following: cameras, LiDAR, radar, maps, navigation, and ego-motion signals.
- Strong understanding of 3D geometry, coordinate frames, calibration, temporal synchronization, ego-motion compensation, tracking, uncertainty estimation, and sensor failure modes.
- Experience with deep-learning methods for 3D perception, scene representation, occupancy and occlusion prediction, semantic segmentation, object detection and tracking, motion prediction, or planning.
- Strong C++ and Python programming skills, with hands-on experience developing, training, and optimizing deep-learning models in PyTorch.
- Experience with CUDA, distributed training, mixed-precision techniques, and efficient GPU inference using NVIDIA software and hardware is highly valued.
- Experience with Transformer, VLM, or multimodal foundation-model architectures, including pre-training, fine-tuning, distillation, quantization, or efficient inference.
- Experience training and evaluating models at scale, including distributed training, dataset curation, offline evaluation, simulation-based validation, and production monitoring.
- Ability to work across research and engineering boundaries by turning ambiguous autonomous-driving problems into measurable technical objectives, building solutions, and driving them to deployment.
Preferred Qualifications
- Experience building multi-task driving models that jointly predict perception, road structure, occupancy, motion, and/or trajectories from shared multimodal features.
- Experience with BEV, point-cloud/voxel, neural scene representation, 3D reconstruction, occupancy-flow, or spatiotemporal world-model methods.
- Publications or open-source contributions in computer vision, robotics, machine learning, 3D perception, multimodal learning, or autonomous driving.
- Experience optimizing models for automotive-grade real-time deployment using NVIDIA GPUs, TensorRT, CUDA, or edge inference toolchains.
Compensation and Benefits
- Level 4 base salary: USD 184,000–287,500 per year.
- Level 5 base salary: USD 224,000–356,500 per year.
- Base salary is determined based on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until September 22, 2026.
- NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior Applications Software Engineer, Perception and Sensor Fusion
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Engineer, Security Architecture - DGX Cloud
Nvidia · United States
USD 272,000-431,200 per year
Senior System Software Engineer – Linux-Tegra Power Management and Performance Optimization
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
PhD Research Intern, Learning Embodied Skills from Human Data - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Senior Synthetic Data Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Integration Engineer, End-to-End Model - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Systems Software Engineer, Semiconductor Systems Inspection
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior System Software Engineer, Interactive World Models
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Metropolis Vision AI
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NVIDIA 2027 Internships: Software Engineering
Nvidia · Santa Clara, United States
USD 20-71 per hour