Tencent Hunyuan has released an open-source solution for those who want to run LLMs locally on a coffee machine.
🟢 How it actually done?
HY-1.8B-2Bit - a model that has been compressed so tightly that it takes up less space than many modern mobile applications.
The model was trained using Quantization-Aware Training, which, unlike PTQ, allows adaptation to low-precision weights even at the training stage.
They took the backbone Hunyuan-1.8B-Instruct and tightly compressed the weights to 2 bits. At the same time, the effective size in memory turned out to be equivalent to a model with 300M parameters, and the physical size was only 600 MB.
Most importantly, they preserved the Dual-CoT feature: the model can switch between quick thinking for simple tasks and deep long-CoT for complex ones.
➜ BENCHMARKS
☞ Compared to the fp16 teacher (1.8B), the degradation of metrics is only ~4%. This is very little for 2-bit quantization.
☞ Difference in accuracy compared to INT4 is negligible - 0.13%, although the model weighs 2 times less.
☞ If you take a dense model with 0.5B parameters, then HY-1.8B-2Bit outperforms it on average by 16-17%. On GSM8K, the gap is even wilder: +22.29%.
☞ Prefill has accelerated by 3-8 times, and token generation by 2-3 times on supported hardware.
➜ A CRUCIAL NUANCE
The current implementation requires support for Arm SME2 instructions. This means that all this beauty will only work on Apple M4 and MediaTek Dimensity 9500.
If you have M1/M2 or previous-generation Snapdragon - it's not for you yet. The developers promise to add the Neon kernel later.
By the way, GGUF is also available, so if you have an M4 - you can test it. The rest have to wait for optimization for old instructions.
Model, GGUF, Technical report, GitHub | #AI #ML #SLM #2bitQ #Tencent
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore