Senior System Software Engineer, Agentic Inference - Dynamo

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ Hybrid

Tech Stack

AI API @ 4 Agile @ 6 Claude Code @ 4 Codex @ 4 Debugging @ 6 Deep Learning GPU @ 4 GenAI Generative AI LLM @ 4 Networking @ 4 Performance Analysis @ 6 Python @ 6 Rust @ 6 SGLang TensorRT vLLM

Details

What you'll be doing

  • Develop open source software to serve inference of trained AI models running on GPUs.
  • Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads, including long-horizon reasoning, tool calling, and stateful, multi-turn execution.
  • Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL, to reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted LLMs.
  • Build and evolve Dynamo's distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, delivering day-0 support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics.
  • Balance a variety of objectives: build robust, scalable, high performance software components to support distributed inference workloads; work with team leads to prioritize features and capabilities; load-balance asynchronous requests across available resources; optimize throughput under latency constraints; and integrate the latest open source technology.

Requirements

  • Masters or PhD or equivalent experience
  • 10+ years in Computer Science, Computer Engineering, or related field
  • Ability to work in a fast-paced, agile team environment
  • Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design.
  • Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs.

Ways to stand out from the crowd

  • Prior contributions to open-source AI inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang).
  • Experience optimizing GPU memory, KV and prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads.
  • Understanding of LLM-specific inference challenges for agentic workloads, including context and reasoning-token growth, bursty tool-call-driven traffic, multi-turn state reuse, and scheduling across concurrent trajectories.
  • Prior experience integrating self-hosted LLM serving stacks with agent harnesses such as OpenCode, Codex, Claude Code, and Pi, including compatibility for APIs, streaming, structured outputs, tool calls, and session semantics.

Additional information

  • NVIDIA is hiring software engineers for its GPU-accelerated deep learning software team, building a Generative AI inference platform.
  • Base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.
  • Applications accepted at least until July 31, 2026.

Location / work style

#LI-Hybrid

More jobs at Nvidia

Similar jobs