Senior Software Engineer, AI Inference

at Nvidia
📍 Toronto, Canada
CAD 135,000-220,000 per year
SENIOR
✅ Hybrid

Tech Stack

AI Communication @ 7 Debugging @ 4 GPU @ 7 Kubernetes @ 6 LLM @ 4 Leadership @ 7 Machine Learning @ 6 Performance Analysis @ 3 Profiling @ 7 SGLang @ 6 Slurm @ 6 vLLM @ 4

Details

Help push the boundaries of AI inference at NVIDIA, where systems expertise shapes both the technology and the teams building on top of it.

The role focuses on large-scale LLM serving and involves partnering with technically demanding customers to unlock the performance potential of NVIDIA's inference stack. Responsibilities combine deep systems knowledge with hands-on customer engagement, including profiling real deployments, benchmarking across GPU clusters, and contributing improvements to the open-source ecosystem. The role includes contributions to vLLM, NVIDIA Dynamo, and productivity tooling.

Responsibilities

  • Work directly with customer engineering teams through long-term technical partnerships to understand LLM serving architectures and performance goals.
  • Design and implement end-to-end benchmarking campaigns across Kubernetes and Slurm environments.
  • Set up and operate vLLM serving deployments on GPU clusters, tuning configurations for throughput, latency, and efficiency.
  • Collect Nsight Systems and Nsight Compute profiling traces to identify performance gaps relative to reference frameworks.
  • Develop performance plans based on profiling findings.
  • Collaborate with NVIDIA kernel engineering and open-source vLLM teams to drive improvements.
  • Build internal tools, benchmarking harnesses, and automation pipelines.
  • Document architectures, findings, and recommendations for technical audiences.
  • Contribute improvements to vLLM and related open-source projects where appropriate.

Requirements

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 5+ years of industry experience building and operating complex, production-grade software systems.
  • Hands-on experience deploying and operating LLM inference workloads, particularly with vLLM, including configuration, optimization, and debugging in real-world environments.
  • Proficiency with Kubernetes and Slurm for running GPU-accelerated workloads.
  • Understanding of LLM serving fundamentals, including continuous batching, chunked prefill, KV cache management, and tensor and pipeline parallelism.
  • Familiarity with GPU performance analysis, including memory hierarchy, utilization, roofline modeling, and profiling with Nsight Systems or Nsight Compute.
  • Strong written and verbal communication skills, with the ability to present technical findings to engineering teams and leadership and navigate ambiguous customer problems.

Preferred Qualifications

  • Experience with NVIDIA Dynamo or other disaggregated inference serving frameworks.
  • Contributions to open-source inference or machine learning systems projects, particularly vLLM or SGLang.
  • Background with ML compilers or GPU kernel development, including Triton, CUTLASS, or TorchInductor.
  • Experience building developer tools or internal platforms that improved team productivity.
  • Prior experience in a customer-facing or forward-deployed engineering capacity within a technical product organization.

Benefits

NVIDIA offers competitive salaries, a comprehensive benefits package, equity, and benefits for employees and their families. The posting indicates that applications will be accepted at least until April 14, 2026 and is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs