Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 4
Computer Vision @ 7
Data Pipelines @ 4
Deep Learning @ 8
GPU @ 4
JAX @ 8
Linux @ 8
PyTorch @ 8
Python @ 8
Reinforcement Learning @ 4
Robotics @ 4
Slurm @ 4
Spark @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Cosmos is an open platform of world foundation models for Physical AI, built to interpret images, video, and text and turn them into a structured understanding of physical scenes, including motion, object interactions, geometry, and physical context. These models support robots, autonomous vehicles, and smart infrastructure.
The Cosmos Engineering team is hiring a Senior Deep Learning Engineer to own the data engine and end-to-end training verification behind 3D spatial reasoning and perception capabilities, including 2D and 3D grounding, metric geometry, spatial reference frames, cross-view correspondence, and embodied spatial reasoning. The role involves determining what models learn geometry from, verifying that they learned it, and working with Cosmos Lab research scientists to turn 3D research hypotheses into measurable capabilities in shipped models.
Responsibilities
- Own the 3D data engine for Cosmos spatial reasoning by sourcing, curating, filtering, and balancing large-scale real-world image and video corpora into vision-language training data.
- Build annotation and auto-labeling pipelines that produce 3D-grounded supervision at scale, including camera-relative 3D boxes, referring and spatial question answering, free space and reachability, ego-, world-, and object-centric reference frames, cross-view correspondence, camera motion, distance and size, and chain-of-thought traces.
- Validate annotations using programmatic and model-based critics.
- Own data quality end to end, including semantic deduplication, automated quality scoring for faithfulness, completeness, and correctness, coverage analysis, and sampling strategies for balanced pre-training and supervised fine-tuning mixtures.
- Verify end-to-end model training by running and validating pre-training and supervised fine-tuning pipelines, maintaining reproducibility, identifying data and checkpoint regressions, diagnosing throughput and loss anomalies, and attributing capability changes to specific data and recipe decisions.
- Build and operate 3D and spatial evaluation suites using public benchmarks such as CV-Bench, BLINK, RefSpatial, VSI-Bench, SPAR-Bench, and RoboSpatial, as well as NVIDIA's VANTAGE-Bench and internally designed benchmarks.
- Maintain continuous evaluation and traceability from reported scores to the exact weights, inputs, configuration, and evaluation code.
- Partner with Cosmos Lab research scientists to translate 3D research hypotheses into dataset and ablation experiments, run experiments at scale, and inform recipe and architecture decisions.
- Operate on large multi-node GPU clusters and tune data throughput, sharding, and dataloader performance.
- Ship results into Cosmos releases and open-source datasets and benchmarks where appropriate.
Requirements
- MS or PhD in Computer Science, Electrical or Computer Engineering, Robotics, or a related field, or equivalent experience.
- 12+ years of experience building deep learning systems in Python with PyTorch or JAX on Linux.
- Deep expertise in 3D computer vision, multi-view geometry, structure-from-motion or SLAM, depth and camera pose estimation, point cloud processing, or 3D reconstruction, with practical experience producing and validating 3D ground truth at scale.
- Hands-on experience with vision-language models, including building training data and evaluations that improve visual grounding and reasoning quality.
- Experience building large-scale multimodal data pipelines for distributed video and image processing, deduplication, captioning and annotation, automated quality metrics, and dataset versioning.
- Experience running and validating large model training on multi-GPU, multi-node clusters, with knowledge of distributed training and sharding strategies such as data, tensor, and pipeline parallelism or FSDP.
- Rigorous evaluation methodology, including designing benchmarks resistant to gaming, building clean ablations, and analyzing results objectively.
- Excellent written and verbal communication, with experience partnering with research scientists and translating research direction into engineering execution.
Preferred Qualifications
- PhD and/or publications at CVPR, ICCV, ECCV, NeurIPS, ICLR, or CoRL in 3D vision, multimodal learning, or embodied AI.
- Experience curating web-scale or petabyte-scale video corpora and building distributed processing infrastructure using Ray, Spark, Slurm, or similar technologies.
- Familiarity with 3D foundation models and modern auto-labeling techniques for geometry, camera estimation, and correspondence.
- Experience with vision-language model post-training, supervised fine-tuning, chain-of-thought data design, reward modeling, or reinforcement learning with verifiable rewards applied to reasoning quality.
- Experience building evaluation infrastructure and harnesses such as VLMEvalKit, including leaderboards, dashboards, and example-level failure inspection.
Benefits
The position offers a base salary range of USD 224,000 to USD 356,500, determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.