Senior Software Engineer, Quantized Inference

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 7 Communication @ 7 Data Analysis Debugging @ 6 LLM @ 4 Machine Learning PyTorch @ 3 Python @ 3 SGLang @ 4 vLLM @ 4

Details

NVIDIA is seeking a Senior Software Engineer to accelerate the discovery and deployment of efficient inference recipes for large language models. A recipe defines which operators are transformed into low-precision or sparsified variants to improve throughput and latency without regressing accuracy or verbosity. Recipes may use rotations, block scaling to attenuate outlier impact, or improved calibration data from SFT/RL pipelines.

Each recipe requires corresponding kernel and model-level implementations in inference engines such as vLLM, TRT-LLM, and SGLang. The role involves translating recipe specifications into functionally correct, performant code, including writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in mixture-of-experts layers is handled correctly. The engineer will collaborate with partner inference teams to optimize throughput and interactivity on target workloads. This work supports productization across Megatron-LM, ModelOpt, and vLLM.

Responsibilities

  • Implement quantized and sparse recipes in vLLM, TRT-LLM, and SGLang.
  • Own model export pipelines between ModelOpt, Megatron-LM, and Hugging Face, ensuring quantized checkpoints serialize correctly for downstream serving.
  • Build prototypes and benchmarking harnesses to evaluate recipe throughput and interactivity before full optimization.
  • Develop data analysis tooling and visualizations for numerics debugging.
  • Improve developer productivity across CI, build systems, training infrastructure, and pipeline workflows.
  • Participate in code reviews and incorporate feedback.

Requirements

  • Proficiency in Python and familiarity with C++.
  • Strong software engineering fundamentals, including concise, well-tested code; fluency with AI-assisted tooling.
  • Experience with ML accelerators and a basic understanding of how ML layers affect execution time.
  • Familiarity with PyTorch internals, including custom operators, autograd, and export, or an equivalent framework.
  • Experience reading, modifying, or contributing to a large open-source codebase.
  • MS or PhD in Computer Science or a related field, or equivalent experience.
  • Four or more years of experience in a relevant software engineering role.
  • Ability to move quickly with ambiguous requirements, along with strong written and verbal communication skills.

Preferred Qualifications

  • Experience contributing to inference serving frameworks such as vLLM, TRT-LLM, or SGLang, or developing Triton kernels.
  • Track record of debugging numerical issues across mixed-precision boundaries.
  • Deep experience with model compression techniques, including post-training quantization (PTQ), quantization-aware training (QAT), and structured or unstructured sparsity.

Compensation and Benefits

  • Base salary range: USD 152,000–241,500 for Level 3.
  • Base salary range: USD 184,000–287,500 for Level 4.
  • Salary is determined based on location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until July 26, 2026.
  • NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs