Senior System Software Engineer - Dynamo-Triton Inference Server
at Nvidia
USD 224,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile @ 7
Communication @ 7
Debugging @ 7
Deep Learning @ 8
Distributed Systems @ 4
GPU @ 4
GitHub @ 4
LLM @ 4
Machine Learning @ 4
NLP
Networking @ 4
Performance Analysis @ 7
PyTorch @ 4
Python @ 3
Rust @ 6
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring a Senior System Software Engineer for the GPU-accelerated deep learning software team working on the Dynamo-Triton Inference Server. The team is building a high-performance AI inference platform that makes the design and deployment of new AI models easier and more accessible. The platform supports workloads ranging from image classification and speech recognition to natural language processing.
Responsibilities
- Develop GPU-accelerated AI inference serving software.
- Contribute to feature development and drive broad customer adoption.
- Drive the convergence of the Triton Inference Server and NVIDIA Dynamo stacks to establish a unified, high-performance inference platform with feature parity for Large Language Model (LLM) and non-LLM workloads.
- Participate actively in the open-source deep learning software engineering community.
- Build robust software for production server and cloud environments.
- Optimize and balance prediction throughput and latency.
- Develop and adopt next-generation inference technologies.
Requirements
- Master's or PhD degree in Computer Science or a relevant field, or equivalent experience.
- 12+ years of professional experience working on deep learning software.
- Excellent Rust and C++ skills.
- Familiarity with Python.
- Strong programming and software design skills, including debugging, performance analysis, and test design.
- Experience with high-scale distributed systems and machine learning systems.
- Strong communication skills and the ability to work in a fast-paced, agile team environment.
Preferred Qualifications
- Experience with AI frameworks and engines such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM.
- Knowledge of GPU memory management, cache management, or high-performance networking.
- Experience with distributed systems programming.
- Experience contributing to a large open-source project, including GitHub, bug tracking, branching and merging code, open-source licensing issues, and handling patches.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications for this job will be accepted at least until September 22, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Software Engineer, Capacity Management - DGX Cloud
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Poland
PLN 221,200-507,000 per year
Tech Lead - Cryptographic Asset Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NVIS Strategy Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Compiler Engineer, Agentic Compiler Systems
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior System Software Engineer - Dynamo-Triton Inference Server
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Software Engineering Intern, Dynamo – Fall 2026
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior System Software Engineer, Agentic Inference – Dynamo
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year