Most RAG systems simply burn the budget. They pull out 100 chunks when you actually need 10. They force LLMs to digest thousands of irrelevant tokens. In the end, you pay for computations that aren't even needed.
🟢 How Meta AI has solved this problem?
They created REFRAG, a new approach to RAG that compresses and filters the context before it even reaches the LLM.
The results sound extremely intriguing:
• 30.85 times faster time-to-first-token
• 16 times larger context windows
• 2-4 times fewer processed tokens
• outperforms LLaMA on 16 RAG benchmarks
What REFRAG does differently: classic RAG simply dumps everything into the LLM. Every chunk. Every token. Even the irrelevant ones.
REFRAG works at the level of embeddings:
↳ compresses each chunk into a single embedding
↳ an RL policy (trained via reinforcement learning) scores each chunk for relevance
↳ only the best chunks are expanded and sent to the LLM
↳ the rest remains compressed or is filtered out entirely
In other words, the LLM only processes what matters.
The pipeline is simple:
1. Encode documents and store them in a vector database
2. When a query arrives, retrieve the relevant chunks as usual
3. The RL policy evaluates the compressed embeddings and selects the best ones
4. The selected chunks are expanded into full token embeddings
5. The rejected chunks remain as single compressed vectors
6. Everything together goes to the LLM
The result: you can process 16 times more context 30 times faster without losing accuracy.
Paper 📝
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
