Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Agentic AI
Agentic Systems @ 3
CUDA @ 3
Communication @ 3
Computer Vision
Deep Learning @ 7
LLM
Machine Learning @ 7
Parallel Programming @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an outstanding PhD intern to work on efficient deep learning with the Deep Learning Efficiency Research (DLER) team. The team focuses on research with real-world impact in two core areas: efficient diffusion language models and multimodal generative models, and efficient agentic AI with hybrid inference orchestration across cloud and edge. Additional areas of interest include post-training model optimization, pruning, quantization, neural architecture search (NAS), efficient architecture design, adaptive and dynamic inference, and resource-efficient training and fine-tuning.
The team has expertise in computer vision, deep learning, generative models, diffusion LLMs, multimodal models, and hybrid cloud-edge agentic systems. Interns will have opportunities to contribute to products, publish research, and collaborate with internal and external researchers.
Responsibilities
- Research, design, and implement novel methods for efficient deep learning in one or both of the following areas:
- Diffusion LLMs and multimodal models: sampling efficiency, adaptive unmasking, self-speculation and parallel decoding, training and distillation pipelines, and multimodal generation.
- Efficient agentic AI: hybrid inference orchestration across cloud and edge, routing and scheduling policies, on-device versus cloud expert delegation, and resource-aware agent loops.
- Publish original research.
- Collaborate with team members and other teams.
- Work with product groups to transfer technology.
- Collaborate with external researchers.
Requirements
- Pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Excellent knowledge of machine learning and deep learning theory and practice.
- Required experience with large language models, diffusion language models, multimodal or vision-language models, or agentic systems.
- Required hands-on experience with large-scale model training, including data preparation and tensor and pipeline model parallelization.
- Outstanding research track record, including at least one publication at a top-tier conference such as ICML, ICLR, NeurIPS, CVPR, or ICCV.
- Excellent communication skills.
Preferred Qualifications
- Parallel programming experience, such as CUDA.
- Interest or experience in hybrid cloud-edge inference, orchestration, or adaptive routing.
- Background in pruning, quantization, NAS, or efficient backbones.
Benefits
- Intern benefits are available.
- NVIDIA provides competitive salaries and a benefits package.
Applications will be accepted at least until October 9, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.