🟢 What is the smart block quantization idea?
- All values are divided into blocks of 16 numbers.
- Each block has its own local scale (8 bits).
- The entire tensor gets a global scale (32 bits).
This preserves high precision of local values and does not lose extremely large or small numbers.
🟣 Results:
- Training the 12B Mamba Transformer on 10T tokens in 4 bits showed accuracy comparable to FP8.
- Computations became 2–3 times faster, and memory usage decreased by 50%.
- Accuracy loss does not exceed 1–1.5% by metrics.
- MMLU Pro: 62.58% (NVFP4) versus 62.62% (FP8).
- MBPP+: 55.91% versus 59.11%.
- Gradients use stochastic rounding to avoid error accumulation.
- Compared to MXFP4, NVFP4 requires 36% less data for the same loss level.
At later training stages, switching to BF16 almost eliminates the quality gap.
NVFP4 is already supported in the Transformer Engine and on Blackwell GPU, including all necessary rounding modes.
🤖 Data Science, ML & Big Data with @DataXplore
