Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms
Debugging @ 6
Deep Learning
Distributed Systems @ 6
LLM
Leadership @ 6
Machine Learning @ 6
Networking
Profiling @ 6
Python @ 7
Rust @ 7
SGLang
TensorRT
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Deep Learning Algorithms Engineer to advance Dynamo, its open-source distributed inference platform for large-scale, low-latency AI services. You will lead architecture and performance work across Dynamo and open-source frameworks, collaborating with research, software, systems, and hardware teams to make AI inference faster, more efficient, and easier to deploy. You will engage with the broader ecosystem, including vLLM, SGLang, and TensorRT-LLM, as well as external partners.
Responsibilities
- Design, build, and maintain Dynamo integrations for the open-source frameworks vLLM, SGLang, and TensorRT-LLM.
- Partner with open-source communities to achieve measurable gains in latency, throughput, reliability, and efficiency.
- Showcase NVIDIA token-per-watt leadership by pushing the Pareto frontier on public and private benchmarks.
- Find and remove bottlenecks across runtimes, kernels, networking, routing, and orchestration.
- Develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent experience.
- At least 3 years of experience building, profiling, and debugging performance-critical distributed or machine learning systems.
- Strong programming skills in Python and/or Rust and C++.
- Understanding of modern machine learning architectures and inference techniques.
Preferred Qualifications
- High agency and a track record of leading ambiguous work end to end.
- Experience with AI accelerators.
- Open-source contributions or leadership.
- Research experience in machine learning inference or distributed systems.
Benefits
The role includes eligibility for equity and NVIDIA benefits. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior System Software Engineer - Dynamo-Triton Inference Server
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Software Engineering Intern, Dynamo – Fall 2026
Nvidia · Santa Clara, United States
USD 20-71 per hour