Senior Applied Research Scientist, Data Curation

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Communication @ 6 Computer Vision @ 4 Data Pipelines Deep Learning @ 4 GPU GitHub @ 6 HTML LLM Machine Learning @ 4 Mentoring @ 1 Microservices PyTorch @ 4 Python @ 4 Spark @ 6

Details

NVIDIA’s Curator team is seeking a Senior Applied Research Scientist to research, develop, and deploy deep learning models at scale across multiple modalities. The role focuses on data curation and document extraction pipelines for foundation-model training, including structured content extraction, data quality, and petabyte-scale deduplication. The work supports NVIDIA Nemotron Parse and the training of Nemotron large language models.

Responsibilities

  • Develop efficient, performant models and data pipelines for extracting and curating multimodal data, including documents, images, audio, and video.
  • Build petabyte-scale extraction and content-deduplication pipelines, including document and HTML parsing, fuzzy and near-duplicate deduplication, semantic deduplication, and substring deduplication.
  • Expand and optimize data-curation methodologies for multimodal datasets processed across hundred-node GPU clusters.
  • Design datasets, metrics, experiments, and validation scripts for standardized research methodologies.
  • Help scale pipelines for production through NVIDIA Inference Microservices (NIMs) and deployment blueprints.
  • Write papers, blog posts, documentation, and training materials to communicate research findings.
  • Keep current with developments in data curation across academia and industry.

Requirements

  • Master’s degree, Ph.D., or equivalent experience in data curation, document AI, information retrieval, multimodal research, or a related field.
  • Track record of publications at leading conferences such as CVPR, ICCV, ECCV, or KDD.
  • Hands-on experience developing computer vision and document-extraction models and pipelines, including layout analysis, OCR, and table, figure, or formula extraction.
  • Understanding of data-curation research, particularly multimodal content extraction and deduplication.
  • 10+ years of experience developing multimodal systems across a range of models and platforms.
  • Experience with information retrieval is a plus.
  • Expertise managing distributed data frameworks such as Ray, Spark, or Dask.
  • Experience deploying massive, multi-node machine learning or data-processing tasks in production environments.
  • Knowledge of batching, streaming, and ingestion-pipeline scaling best practices.
  • Excellent Python programming skills and hands-on experience with PyTorch or comparable modern deep-learning frameworks.
  • Ability to communicate ideas clearly through blog posts, papers, kernels, GitHub, and similar formats.
  • Excellent communication and interpersonal skills, with the ability to work in a dynamic, user-focused, distributed team.
  • Experience mentoring junior engineers and interns is a plus.
  • Kaggle Grandmaster status or a strong record in machine-learning competitions is a plus.

Benefits

  • Base salary range of USD 224,000–356,500 per year, determined by location, experience, and comparable employee compensation.
  • Eligibility for equity and benefits.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

The team is remotely situated and focuses on North American and European time zones. Remote work is accepted for candidates in any country where NVIDIA has an office. Applications will be accepted at least until August 31, 2026.

More jobs at Nvidia

Similar jobs