TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #2033 179
OVQA: goodbye, KV-cache offloading.

Zyphra came up with a way to sit on two chairs at once when you want rubbery context but don't have a ton of memory on hand.

What they proposed is called Online Vector-Quantized Attention - it's a modification of vector quantization that teaches the dictionary to think on the fly.

🟢 What they done?

In classic VQ, keys are replaced by the nearest centroids from a static dictionary. This boosts computations but creates a problem: the dictionary is trained on one set of data, but during generation, the model sees a completely different distribution of keys. The quantization error grows, attention loses accuracy, and as a result: VQ starts to falter.

So, the modification is to abandon the static dictionary in favor of one adapted to the current sequence: each new token updates only one centroid - the one it's closest to.

This sparse update works as a protection against catastrophic forgetting: old information isn't washed away by a new wave of tokens, but is carefully overwritten as needed.

There's also a hard limit on the state size, after which the amount of memory stops growing and computations become strictly linear.

➜ Results of test experiments

☞ A model trained on 4K tokens confidently handled context up to 64K without quality degradation;

☞ On in-context search, OVQ almost kept up with full self-attention while consuming 4 times less memory;

☞ On In-Context Learning, VQ failed, but OVQ reached the level of classic attention using only ~4K centroids;

☞ Comparisons with linear alternatives (Mamba2 and delta networks) also favor OVQ: it more stably holds long context without accuracy drops;

➜ In Positional ICR tasks, OVQA performs slightly worse than classic attention but still decently.

We really hope that OVQ is the precursor of true continuous learning, where in the bright future, instead of an infinitely swelling KV-cache, there will be compact but living memory capable of holding important details without losses.


Article, PDF • #AI #ML #LLM #OVQA #Zyphra

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →