A method (by microsoft) is proposed that works for any pair of models, even from different families, different companies and different architectures.
🟢 How models can communicate in their "own language"?
When two agents communicate in a multimodal system, they usually do so via text. This is quite inefficient because each model actually has a Key-Value Cache, an internal attention states that essentially store all the information about the model's thoughts. And if agents could learn to communicate not with tokens but with the KV cache itself, it would be much faster and the information would be more complete.
Thus appears Cache-to-Cache (C2C) which is paradigm of direct exchange of meaning, not words. The source (Sharer) sends its cache, and the receiver (Receiver) embeds this cache into its own space through a neural network projector.
Directly, without a projector, this would not be possible because different models have different hidden spaces. Therefore, the authors trained a Projection module that connects the caches of the Sharer and Receiver into a single embedding understandable to both models. Besides the Projection module, the protocol also includes a weighting module that decides what information from the Sharer is worth transmitting.
THIS GIVES…
1️⃣ Speed, obviously. Compared to Text-to-Text, everything happens 2-3 times faster.
2️⃣ Accuracy improvement : If two models are combined this way and tasked with solving one problem, the metric improves on average by 5% compared to the case when models are combined but communicate via text.
A big practical downside is that the approach is not universal. For each pair of models, you have to train your own "bridge." There are only a few MLP layers, but still. Also, if the models have completely different tokenizers, it’s a hassle and you’ll have to do Token alignment.
By exchanging caches, models really understand each other better than when exchanging tokens. This is a cool result.
GitHub
🤖 Data Science, ML & Big Data with @DataXplore
