tonbistudio/turboquant-pytorch
From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.
Language: Python
Stars: 586 Issues: 12 Forks: 78
https://github.com/tonbistudio/turboquant-pytorch
Post #13103
1.65K