Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 6
Computer Vision @ 4
Data Pipelines
Deep Learning @ 4
GPU
GitHub @ 6
HTML
LLM
Machine Learning @ 4
Mentoring @ 1
Microservices
PyTorch @ 4
Python @ 4
Spark @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s Curator team is seeking a Senior Applied Research Scientist to research, develop, and deploy deep learning models at scale across multiple modalities. The role focuses on data curation and document extraction pipelines for foundation-model training, including structured content extraction, data quality, and petabyte-scale deduplication. The work supports NVIDIA Nemotron Parse and the training of Nemotron large language models.
Responsibilities
- Develop efficient, performant models and data pipelines for extracting and curating multimodal data, including documents, images, audio, and video.
- Build petabyte-scale extraction and content-deduplication pipelines, including document and HTML parsing, fuzzy and near-duplicate deduplication, semantic deduplication, and substring deduplication.
- Expand and optimize data-curation methodologies for multimodal datasets processed across hundred-node GPU clusters.
- Design datasets, metrics, experiments, and validation scripts for standardized research methodologies.
- Help scale pipelines for production through NVIDIA Inference Microservices (NIMs) and deployment blueprints.
- Write papers, blog posts, documentation, and training materials to communicate research findings.
- Keep current with developments in data curation across academia and industry.
Requirements
- Master’s degree, Ph.D., or equivalent experience in data curation, document AI, information retrieval, multimodal research, or a related field.
- Track record of publications at leading conferences such as CVPR, ICCV, ECCV, or KDD.
- Hands-on experience developing computer vision and document-extraction models and pipelines, including layout analysis, OCR, and table, figure, or formula extraction.
- Understanding of data-curation research, particularly multimodal content extraction and deduplication.
- 10+ years of experience developing multimodal systems across a range of models and platforms.
- Experience with information retrieval is a plus.
- Expertise managing distributed data frameworks such as Ray, Spark, or Dask.
- Experience deploying massive, multi-node machine learning or data-processing tasks in production environments.
- Knowledge of batching, streaming, and ingestion-pipeline scaling best practices.
- Excellent Python programming skills and hands-on experience with PyTorch or comparable modern deep-learning frameworks.
- Ability to communicate ideas clearly through blog posts, papers, kernels, GitHub, and similar formats.
- Excellent communication and interpersonal skills, with the ability to work in a dynamic, user-focused, distributed team.
- Experience mentoring junior engineers and interns is a plus.
- Kaggle Grandmaster status or a strong record in machine-learning competitions is a plus.
Benefits
- Base salary range of USD 224,000–356,500 per year, determined by location, experience, and comparable employee compensation.
- Eligibility for equity and benefits.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
The team is remotely situated and focuses on North American and European time zones. Remote work is accepted for candidates in any country where NVIDIA has an office. Applications will be accepted at least until August 31, 2026.
More jobs at Nvidia
Senior Systems Software Engineer, Low Latency Streaming Technology - Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Reinforcement Learning Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Manager, Storage Engineering
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Software and System Architect
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Customer Technical Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Staff Machine Learning Engineer, Consumer
Reddit · United States
USD 230,000-322,000 per year
Senior System Software Engineer - Dynamo-Triton Inference Server
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year