Principal Research Scientist, Synthetic Data Generation

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 7 CI/CD Data Pipelines @ 4 Git LLM Machine Learning @ 4 Statistics @ 4 vLLM @ 4

Details

NVIDIA is seeking a Principal Research Scientist to set the technical direction for synthetic data generation across its frontier model efforts. The role involves defining and building open-source libraries within the NVIDIA NeMo ecosystem to generate synthetic datasets across text, code, structured, and multimodal data for pre-training and post-training large language models such as Nemotron. This position combines hands-on software engineering with applied research in generative methods and involves collaboration with research, engineering, product, model teams, and external labs.

Responsibilities

  • Build and scale data-generation pipelines using LLM-based methods and automated quality evaluation for datasets supporting the initial training and fine-tuning of LLMs such as Nemotron.
  • Develop pipelines covering reasoning, coding, structured output, and multimodal understanding.
  • Pioneer data generation for agentic and tool-use training, including synthetic trajectories, multi-turn interactions, function calling, executable reinforcement-learning environments, reward modeling, and verifiable-reward data.
  • Advance multimodal synthetic data generation for images, documents, video, and audio in partnership with NVIDIA model teams.
  • Advance privacy-preserving and safe synthesis using differential privacy, anonymization, and de-identification for model training on sensitive data in regulated domains.
  • Develop and maintain open-source libraries and SDKs with clean APIs and strong documentation.
  • Drive software excellence through modern tooling, configuration-based architecture, and professional Git and CI/CD practices.
  • Publish original research at leading machine learning and AI conferences.
  • Mentor scientists and engineers across the team.

Requirements

  • PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
  • 15+ years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related areas.
  • Deep technical understanding of LLMs and how data influences pre-training, post-training, and reinforcement-learning stages.
  • Experience with inference frameworks such as vLLM or TGI.
  • Proven track record of developing or maintaining software libraries used by a broad developer community.
  • Experience building and optimizing scalable data pipelines for large-scale model training, including throughput, distributed inference, and cluster-scale cost optimization.
  • Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL, or similar.

Preferred Qualifications

  • Significant open-source contributions in machine learning or data tooling with community adoption.
  • Experience with multimodal generation or understanding, including vision-language, document AI, video, or audio.
  • Experience generating data for agentic, tool-use, or reinforcement-learning post-training, including reinforcement-learning environment design.
  • Background in differential privacy, de-identification, or synthetic data for regulated industries such as healthcare, finance, or government.
  • Experience influencing model-training decisions at frontier scale or partnering directly with pre-training and post-training teams.

Benefits

  • Equity and NVIDIA benefits.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs