Senior Deep Learning Architect, LLM Inference

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 API @ 4 Agentic AI @ 6 Claude Code @ 6 Codex @ 6 Communication @ 6 Deep Learning @ 4 GPU @ 7 GenAI Generative AI LLM @ 4 Marketing OSS Profiling @ 4 PyTorch @ 4 SGLang Software Development @ 4 vLLM

Details

NVIDIA's Inference Benchmarking team focuses on optimizing inference server performance for large language models (LLMs). This role involves pushing the boundaries of GPU hardware and software performance across technologies such as disaggregated serving, data parallel attention, mixture-of-experts (MoE), Qwen3.5, DeepSeek, and GPT-OSS.

Responsibilities

  • Characterize workloads for the latest LLMs and inference servers, including vLLM, SGLang, and TRT-LLM.
  • Collaborate with the performance marketing team to create technical content, including blog posts and InferenceX updates.
  • Work with engineers from AI startup companies to establish standard benchmarking methodologies.
  • Develop and maintain an evolving inference performance data results website.
  • Create end-to-end profiling and analysis tools for generative AI workloads.
  • Contribute to deep learning software projects such as PyTorch, TRT-LLM, vLLM, and SGLang.
  • Verify that new GPU product launches deliver industry-leading performance.
  • Collaborate with software, research, and product teams to guide the direction of inference serving.
  • Use current coding agents and inference technologies to improve team efficiency.

Requirements

  • Master's or PhD degree in Computer Science, Computer Engineering, a related field, or equivalent experience.
  • 6 or more years of relevant software development experience.
  • Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client-server LLM applications using the OpenAI API or MCP and identifying performance bottlenecks.
  • Strong understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with complex software projects such as frameworks, compilers, or operating systems.
  • Proficiency with AI coding agents such as Claude Code, Codex, and Cursor.
  • Excellent written and verbal communication skills, with the ability to work independently and collaboratively in a fast-paced environment.

Preferred Qualifications

  • A demonstrated drive to continuously improve software and hardware performance.
  • Examples of novel workplace use cases for agentic AI tools.
  • Experience with databases and visualization tools.

Compensation

The base salary depends on location, experience, and the pay of employees in similar positions. The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Employees are also eligible for equity and benefits.

Applications will be accepted at least until August 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs