Senior Deep Learning Scientist, Multimodal Agentic RL

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic Systems @ 4 Algorithms @ 4 Deep Learning @ 7 LLM Machine Learning @ 7 Mathematics @ 7 PyTorch @ 7 Python @ 7 Reinforcement Learning @ 4

Details

NVIDIA is hiring Senior Deep Learning Scientists to advance streaming and agentic multimodal AI. You will apply deep learning, reinforcement learning, and applied mathematics to develop models capable of reasoning, planning, and acting across diverse modalities. The role focuses on algorithmic improvements for multimodal foundation models and scaling research through NVIDIA's Nemotron Omni and VoiceChat platforms.

Responsibilities

  • Apply fundamental and applied research to develop, train, fine-tune, and deploy large language models for agentic systems involving audio-visual reasoning, tool usage, and document understanding.
  • Advance post-training and alignment methods, including instruction tuning, preference optimization, RLHF, RLVR, and MOPD, to improve multimodal agents for complex use cases.
  • Research and develop agentic reasoning and grounded perception capabilities, with a focus on planning, tool execution, and long-horizon task completion across digital and physical environments.
  • Lead the collection, development, and benchmarking of multimodal datasets, ensuring high-quality evaluation of model accuracy, safety, and task-completion success.

Requirements

  • Master's degree or equivalent experience, or PhD, in Computer Science, Artificial Intelligence, or Applied Mathematics, with 8+ years of relevant work experience.
  • Excellent programming skills in Python and strong fundamentals in scalable model development and deep learning frameworks such as PyTorch.
  • Strong knowledge of machine learning and deep learning techniques and modern foundation model architectures, including Transformers and mixture-of-experts models.
  • Foundational understanding of reinforcement learning algorithms and implementation, including Markov decision processes, policies, and reward design.
  • Hands-on experience post-training multimodal models for omni-modality audio-visual reasoning, full-duplex voice chat, and human-AI interaction.
  • Proven ability to manage model development life cycles, including dataset versioning, experiment tracking, and evaluation pipelines.

Preferred Qualifications

  • Strong publication record in top-tier AI and machine learning venues such as NeurIPS, ICML, ICLR, or CVPR.
  • Experience training and deploying multimodal foundation models using large-scale distributed infrastructure.
  • Experience applying deep reinforcement learning techniques to train multimodal agents in complex simulation or gaming environments.
  • Background in audio or speech AI, especially audio language models or audio generation.
  • Background building embodied AI systems that integrate multimodal perception with backend action fulfillment and long-horizon planning.

Benefits

The role offers a competitive salary, equity, and a comprehensive benefits package. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until October 6, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs