Merging RAG and CAG. How can an AI engineer use this?
Let's break down how this looks and what additional considerations need to be taken into account.
➡️ Example steps for a CAG + RAG architecture:
👉 DATA PREPARATION
1️⃣ For CAG, we only use sources that change infrequently. In addition to "infrequent changes," it's important to understand which of these sources are most frequently needed for relevant queries. Only after this, we "warm up" the selected data in advance in the KV cache model and cache it in memory. This is done once, and the remaining steps can be repeated many times without recalculating the initial cache.
2️⃣ For RAG, if necessary, we pre-calculate and save vector embeddings in a compatible database so that we can later search for them in step 4. Sometimes, simpler types of data are sufficient for RAG, in which case a regular database would work.
👉 QUERY PATH
Now we can use the prepared data.
3️⃣ We assemble the prompt: the user's query + a system prompt with instructions on how the model should use the cached context and external (retrieved) context.
4️⃣ We build an embedding of the user's query for semantic search through the vector DB and query the context store to retrieve relevant data. If semantics aren't needed, we can go to other sources, such as a real-time database or the web.
5️⃣ We enrich the final prompt with the external context retrieved in step 4.
6️⃣ We return the final response to the user.
👉 A FEW IMPORTANT POINTS
☞ The context window is not infinite. Even if the model has a huge context, the "needle in a haystack" problem still exists. Use context sparingly and cache only what is really needed.
☞ For some cases, certain datasets are super valuable to constantly feed them to the model through the cache. For example, an assistant who is obliged to always comply with a long set of internal rules scattered across several documents.
☞ Although open-source CAG has gained popularity relatively recently, in practice, this has long been possible through prompt caching in the OpenAI and Anthropic APIs. It's easy to quickly put together a prototype there.
☞ Always separate hot and cold sources. We only put cold (infrequently changing) data in the cache, otherwise the data will become outdated, and the application will start living "out of reality."
⚠️ RISKS & LIMITATIONS
☞ Be very careful with what you cache, because this data becomes available to all users' queries.
☞ It's difficult to ensure RBAC for the cache if you don't have a separate model with its own cache for each role.
Have you already tried this combo approach?
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
