Senior Software Engineer, Quantized Inference

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 7 Communication @ 7 Data Analysis Debugging LLM @ 4 Machine Learning PyTorch @ 3 Python @ 3 SGLang @ 4 vLLM @ 4

Details

Responsibilities

  • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
  • Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
  • Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
  • Develop data analysis tooling and visualizations for numerics debugging
  • Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction
  • Participate in code reviews and incorporate feedback

Requirements

  • Proficient in Python; familiarity with C++
  • Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling
  • Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time
  • Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework
  • Experience reading, modifying, or contributing to a large open-source codebase
  • MS/PhD in Computer Science or related field, or equivalent experience
  • 4+ years in a relevant software engineering role
  • Demonstrated ability to move fast with ambiguous requirements, with strong written and verbal communication

Ways to stand out from the crowd

  • Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development
  • Track record of debugging numerical issues across mixed-precision boundaries
  • Deep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsity

More jobs at Nvidia

Similar jobs