Senior System Software Engineer - Dynamo-Triton Inference Server

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 4 Agile @ 7 Communication @ 7 Debugging @ 7 Deep Learning @ 8 Distributed Systems @ 4 GPU @ 4 GitHub @ 4 LLM @ 4 Machine Learning @ 4 NLP Networking @ 4 Performance Analysis @ 7 PyTorch @ 4 Python @ 3 Rust @ 6 TensorRT @ 4 vLLM @ 4

Details

NVIDIA is hiring a Senior System Software Engineer for the GPU-accelerated deep learning software team working on the Dynamo-Triton Inference Server. The team is building a high-performance AI inference platform that makes the design and deployment of new AI models easier and more accessible. The platform supports workloads ranging from image classification and speech recognition to natural language processing.

Responsibilities

  • Develop GPU-accelerated AI inference serving software.
  • Contribute to feature development and drive broad customer adoption.
  • Drive the convergence of the Triton Inference Server and NVIDIA Dynamo stacks to establish a unified, high-performance inference platform with feature parity for Large Language Model (LLM) and non-LLM workloads.
  • Participate actively in the open-source deep learning software engineering community.
  • Build robust software for production server and cloud environments.
  • Optimize and balance prediction throughput and latency.
  • Develop and adopt next-generation inference technologies.

Requirements

  • Master's or PhD degree in Computer Science or a relevant field, or equivalent experience.
  • 12+ years of professional experience working on deep learning software.
  • Excellent Rust and C++ skills.
  • Familiarity with Python.
  • Strong programming and software design skills, including debugging, performance analysis, and test design.
  • Experience with high-scale distributed systems and machine learning systems.
  • Strong communication skills and the ability to work in a fast-paced, agile team environment.

Preferred Qualifications

  • Experience with AI frameworks and engines such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM.
  • Knowledge of GPU memory management, cache management, or high-performance networking.
  • Experience with distributed systems programming.
  • Experience contributing to a large open-source project, including GitHub, bug tracking, branching and merging code, open-source licensing issues, and handling patches.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Applications for this job will be accepted at least until September 22, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs