TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 586 subscribers
Post #1739 162
NVFP4 - a new mathod that trains the 12B Mamba Transformer to store numbers in 4 bits instead of 8 or 16, with almost no loss in training quality and without loss of accuracy

🟢 What is the smart block quantization idea?

- All values are divided into blocks of 16 numbers.
- Each block has its own local scale (8 bits).
- The entire tensor gets a global scale (32 bits).

This preserves high precision of local values and does not lose extremely large or small numbers.


🟣 Results:
- Training the 12B Mamba Transformer on 10T tokens in 4 bits showed accuracy comparable to FP8.
- Computations became 2–3 times faster, and memory usage decreased by 50%.
- Accuracy loss does not exceed 1–1.5% by metrics.
- MMLU Pro: 62.58% (NVFP4) versus 62.62% (FP8).
- MBPP+: 55.91% versus 59.11%.
- Gradients use stochastic rounding to avoid error accumulation.
- Compared to MXFP4, NVFP4 requires 36% less data for the same loss level.


At later training stages, switching to BF16 almost eliminates the quality gap.
NVFP4 is already supported in the Transformer Engine and on Blackwell GPU, including all necessary rounding modes.

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →