Senior Vision Language Model Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic AI @ 6 Algorithms @ 6 Communication @ 4 Debugging @ 4 Deep Learning @ 7 Python @ 4 Robotics @ 4

Details

NVIDIA is seeking a senior vision language model engineer to design and build agentic data and training workflows for autonomous vehicles, robotics, and medical applications. The role focuses on developing scalable dataset search platforms and model training capabilities for physical AI developers.

Responsibilities

  • Partner with researchers to develop and evaluate prototypes of vision-language models (VLMs) and vision-language-action models (VLAs) for video search, video understanding, and related applications.
  • Design and implement agentic data workflows that automate data discovery, labeling, evaluation, and retraining.
  • Build, curate, and maintain high-quality multimodal datasets, including video, sensor data, and language/action traces, for end-to-end physical AI problems such as autonomous driving.
  • Explore and productize new data sources, including simulation and synthetic data.
  • Use agentic AI workflows across the full applied research lifecycle.
  • Collaborate with research, model development, performance, and product teams.
  • Contribute to NVIDIA Cosmos Dataset Search and other core NVIDIA platforms and products.

Requirements

  • PhD with 4+ years of relevant experience, MS with 6+ years, or BS or equivalent experience with 8+ years in Computer Science, Computer Engineering, or a related technical field.
  • Strong background in modern deep learning, including transformer-based architectures, video modeling, multimodal VLM/VLA models, or foundation models.
  • Excellent experience training and deploying deep learning models on real-world datasets, including data preprocessing, distributed training, evaluation, debugging, and iterative improvement.
  • Excellent experience with Python and at least one deep learning framework.
  • Current knowledge of research on image and video search in autonomous vehicles, healthcare, robotics, or related physical AI applications.
  • Fluency with agentic AI workflows across the applied research lifecycle, including prototyping algorithms and search pipelines, benchmarking, and integrating prototypes into production codebases.
  • Clear and effective communication skills, with experience working in a dynamic, product- and research-focused team.

Preferred Qualifications

  • Strong publication record at top-tier conferences such as CVPR, NeurIPS, ICML, or ECCV.
  • Patents in video retrieval or a related field.
  • Strong coding architecture skills demonstrated through contributions to large internal or open-source projects.
  • Experience with robotic systems such as autonomous vehicles or humanoid robotics.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until July 11, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs