Senior System Software Engineer, Agentic Inference – Dynamo

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 6 API @ 4 Agile @ 6 Claude Code @ 4 Codex @ 4 Debugging @ 6 GPU @ 4 GenAI Generative AI LLM @ 4 NLP Networking @ 4 Performance Analysis @ 6 Python @ 6 Rust @ 6 SGLang @ 6 TensorRT @ 6 vLLM @ 6

Details

NVIDIA is hiring a Senior System Software Engineer to work on Dynamo, an open-source, GPU-accelerated generative AI inference platform. The team develops software that makes the design and deployment of AI models easier and more accessible, supporting workloads such as image classification, speech recognition, and natural language processing.

Responsibilities

  • Develop open-source software to serve inference for trained AI models running on GPUs.
  • Contribute to disaggregated serving for Dynamo-supported inference engines, including vLLM, SGLang, and TensorRT-LLM, and expand these capabilities for agentic inference workloads such as long-horizon reasoning, tool calling, and stateful, multi-turn execution.
  • Innovate in inference-state management for long-running agents, including KV-cache and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL.
  • Reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted large language models.
  • Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM.
  • Deliver day-zero support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics.
  • Build robust, scalable, high-performance software components for distributed inference workloads.
  • Work with team leads to prioritize features and capabilities.
  • Load-balance asynchronous requests across available resources and optimize throughput under latency constraints.
  • Integrate the latest open-source technologies.

Requirements

  • Master’s degree, PhD, or equivalent experience.
  • 10+ years of experience in computer science, computer engineering, or a related field.
  • Ability to work in a fast-paced, agile team environment.
  • Excellent Rust and Python programming and software design skills, including debugging, performance analysis, and test design.
  • Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs.

Preferred Qualifications

  • Contributions to open-source AI inference frameworks such as vLLM, TensorRT-LLM, or SGLang.
  • Experience optimizing GPU memory, KV caches, prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads.
  • Understanding of LLM-specific inference challenges for agentic workloads, including context and reasoning-token growth, bursty tool-call-driven traffic, multi-turn state reuse, and scheduling across concurrent trajectories.
  • Experience integrating self-hosted LLM serving stacks with agent harnesses such as OpenCode, Codex, Claude Code, and Pi, including compatibility for APIs, streaming, structured outputs, tool calls, and session semantics.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until July 31, 2026.

More jobs at Nvidia

Similar jobs