Software Engineering Intern, Dynamo – Fall 2026

at Nvidia
USD 20-71 per hour
INTERN
✅ On-site

Tech Stack

AI API @ 3 Algorithms @ 3 Communication @ 6 Data Structures @ 3 Debugging @ 3 GPU GitHub Go Kubernetes @ 3 LLM @ 3 Machine Learning @ 3 NLP Python @ 3 Rust @ 3 SGLang @ 3 Software Development @ 3 TensorRT @ 3 vLLM @ 3

Details

NVIDIA Dynamo is an innovative, open-source platform focused on efficient, scalable inference for large language and reasoning models in distributed GPU environments. The platform uses sophisticated techniques in serving architecture, GPU resource management, and intelligent request handling to achieve high-performance AI inference for demanding applications.

As an Applied AI Research Software Engineering Intern on the Dynamo project, you will work on distributed inference challenges, including the Dynamo Kubernetes serving platform, disaggregated serving, dynamic GPU scheduling, intelligent routing, and distributed KV-cache management.

Responsibilities

  • Collaborate on the design and development of the Dynamo Kubernetes stack.
  • Introduce new features to the Dynamo Python SDK and Dynamo Rust Runtime Core Library.
  • Design, implement, and optimize distributed inference components in Rust and Python.
  • Contribute to disaggregated serving for Dynamo-supported inference engines, including vLLM, SGLang, TensorRT-LLM, llama.cpp, and mistral.rs.
  • Improve intelligent routing and KV-cache management subsystems.
  • Contribute to open-source repositories and participate in code reviews.
  • Assist with issue triage on GitHub, work with the community to address issues, capture feedback, and evolve the framework's APIs and architecture.

Requirements

  • Pursuing a BS, MS, or PhD in Computer Science or an equivalent program area.
  • Excellent Golang, Rust, and/or Python programming and software design skills.
  • Experience with debugging, performance and service health analysis, and test design.
  • Good understanding of algorithms and data structures.
  • Solid knowledge of RESTful APIs.
  • Motivation, dedication, curiosity about new technologies, strong communication, planning, and problem-solving skills.

Preferred Qualifications

  • Understanding of machine learning or natural language processing concepts.
  • Experience with software shipping cycles, including development, deployment, release, and continuous integration.
  • Experience with open-source software development.
  • Experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or similar technologies.
  • Experience building and deploying containers in Kubernetes environments.

Compensation and Benefits

  • Internship hourly rate: USD 20–71, based on position, location, year in school, degree, and experience.
  • Eligible for intern benefits.
  • Applications will be accepted at least until August 9, 2026.
  • NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs