Zyphra came up with a way to sit on two chairs at once when you want rubbery context but don't have a ton of memory on hand.
What they proposed is called Online Vector-Quantized Attention - it's a modification of vector quantization that teaches the dictionary to think on the fly.
🟢 What they done?
In classic VQ, keys are replaced by the nearest centroids from a static dictionary. This boosts computations but creates a problem: the dictionary is trained on one set of data, but during generation, the model sees a completely different distribution of keys. The quantization error grows, attention loses accuracy, and as a result: VQ starts to falter.
So, the modification is to abandon the static dictionary in favor of one adapted to the current sequence: each new token updates only one centroid - the one it's closest to.
This sparse update works as a protection against catastrophic forgetting: old information isn't washed away by a new wave of tokens, but is carefully overwritten as needed.
There's also a hard limit on the state size, after which the amount of memory stops growing and computations become strictly linear.
➜ Results of test experiments
☞ A model trained on 4K tokens confidently handled context up to 64K without quality degradation;
☞ On in-context search, OVQ almost kept up with full self-attention while consuming 4 times less memory;
☞ On In-Context Learning, VQ failed, but OVQ reached the level of classic attention using only ~4K centroids;
☞ Comparisons with linear alternatives (Mamba2 and delta networks) also favor OVQ: it more stably holds long context without accuracy drops;
➜ In Positional ICR tasks, OVQA performs slightly worse than classic attention but still decently.
We really hope that OVQ is the precursor of true continuous learning, where in the bright future, instead of an infinitely swelling KV-cache, there will be compact but living memory capable of holding important details without losses.
Article, PDF • #AI #ML #LLM #OVQA #Zyphra
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
