Developing Nemotron 3.5 Lightning NVFP4 With QAD Using NVIDIA Model Optimizer
NVIDIA, Monday, August 17th, 2026
NVIDIA details the quantization-aware distillation pipeline used to build Nemotron 3.5 Lightning in NVFP4 precision.
NVIDIA published a technical walkthrough of developing Nemotron 3.5 Lightning in NVFP4 precision using quantization-aware distillation with NVIDIA Model Optimizer.
The post starts from the premise that teams customize models to hit specific targets for latency, speed, memory and compute. The open NVIDIA Nemotron family is positioned as a base for that customization.
The QAD pipeline is described step by step so teams can reproduce the approach on their own models. Authors include Seonghee Lee, Jenny Chen, Chenhan Yu, James Shen and Trenton Starkey.