Senior Scientist, Synthetic Data Generation

at Nvidia
USD 168,000-304,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 7 Agentic Systems @ 4 CI/CD Data Pipelines @ 4 Git LLM @ 4 Machine Learning @ 4 Statistics @ 4 vLLM @ 7

Details

NVIDIA is seeking a Senior Scientist to advance synthetic data generation for training frontier models. The role combines hands-on software engineering with applied research in generative methods. You will contribute to open-source libraries within the NVIDIA NeMo ecosystem, generating synthetic datasets across text, code, structured, and multimodal data for the pre- and post-training of large language models such as Nemotron. You will collaborate with research, engineering, product, and model teams, as well as external labs.

Responsibilities

  • Build synthetic data generation pipelines using LLM-based methods and automated quality evaluation for reasoning, coding, structured output, and multimodal understanding.
  • Advance multimodal synthetic data generation for images, documents, video, and audio in partnership with NVIDIA's model teams.
  • Design and maintain open-source libraries and SDKs with clean APIs and strong documentation.
  • Drive software excellence through modern tooling, configuration-based architecture, and professional Git and CI/CD practices.
  • Publish original research at leading machine learning and AI conferences.
  • Mentor interns and junior researchers.

Requirements

  • PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
  • At least 3 years of research experience in synthetic data generation, generative modeling, multimodal machine learning, or related areas; comparable experience is also considered.
  • Deep technical understanding of LLMs, the role of data in pre-training and post-training, and inference frameworks such as vLLM or TGI.
  • Proven experience developing or maintaining software libraries used by a broad developer community.
  • Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL, or similar.

Preferred Qualifications

  • Open-source contributions in machine learning or data tooling.
  • Experience with multimodal generation or understanding, including vision-language, document AI, video, or audio.
  • Experience building and optimizing scalable data pipelines for large-scale model training, including throughput and distributed inference optimization.
  • Experience generating data for agentic systems, tool use, or reinforcement-learning post-training.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until August 10, 2026. The posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs