Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 3
Algorithms @ 3
Communication @ 6
Data Structures @ 3
Debugging @ 3
GPU
GitHub
Go
Kubernetes @ 3
LLM @ 3
Machine Learning @ 3
NLP
Python @ 3
Rust @ 3
SGLang @ 3
Software Development @ 3
TensorRT @ 3
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Dynamo is an innovative, open-source platform focused on efficient, scalable inference for large language and reasoning models in distributed GPU environments. The platform uses sophisticated techniques in serving architecture, GPU resource management, and intelligent request handling to achieve high-performance AI inference for demanding applications.
As an Applied AI Research Software Engineering Intern on the Dynamo project, you will work on distributed inference challenges, including the Dynamo Kubernetes serving platform, disaggregated serving, dynamic GPU scheduling, intelligent routing, and distributed KV-cache management.
Responsibilities
- Collaborate on the design and development of the Dynamo Kubernetes stack.
- Introduce new features to the Dynamo Python SDK and Dynamo Rust Runtime Core Library.
- Design, implement, and optimize distributed inference components in Rust and Python.
- Contribute to disaggregated serving for Dynamo-supported inference engines, including vLLM, SGLang, TensorRT-LLM, llama.cpp, and mistral.rs.
- Improve intelligent routing and KV-cache management subsystems.
- Contribute to open-source repositories and participate in code reviews.
- Assist with issue triage on GitHub, work with the community to address issues, capture feedback, and evolve the framework's APIs and architecture.
Requirements
- Pursuing a BS, MS, or PhD in Computer Science or an equivalent program area.
- Excellent Golang, Rust, and/or Python programming and software design skills.
- Experience with debugging, performance and service health analysis, and test design.
- Good understanding of algorithms and data structures.
- Solid knowledge of RESTful APIs.
- Motivation, dedication, curiosity about new technologies, strong communication, planning, and problem-solving skills.
Preferred Qualifications
- Understanding of machine learning or natural language processing concepts.
- Experience with software shipping cycles, including development, deployment, release, and continuous integration.
- Experience with open-source software development.
- Experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or similar technologies.
- Experience building and deploying containers in Kubernetes environments.
Compensation and Benefits
- Internship hourly rate: USD 20–71, based on position, location, year in school, degree, and experience.
- Eligible for intern benefits.
- Applications will be accepted at least until August 9, 2026.
- NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
Senior System Software Engineer, Agentic Inference – Dynamo
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
ML Solution Architect (Early Talent)
Nebius · United States
USD 102-126 per hour
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year
Senior Software Engineer, Machine Learning Inference
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year