The simple strategy to search through 40 million texts in ~200 ms using only a CPU server, 8GB of RAM, and 45GB of disk space.
If you want to try it out immediately, there's a demo of 40 million texts from Wikipedia. No login or other hassles required.
🟢 The Inference Strategy:
☞ Embed the query with a dense model into a regular fp32 vector
☞ Quantize the fp32 embedding into a binary format, which is 32 times smaller
☞ Retrieve, for example, 40 documents (about 20 times faster than an fp32 index) using an approximate or exact binary index
☞ Load the int8 embeddings for these top-40 documents from the disk
☞ Rescoring: fp32 embedding of the query × 40 int8 embeddings
☞ Sort these 40 documents by the new score and take the top-10
☞ Load the titles and texts of the top-10 documents
The documents are embedded once, and then these embeddings are used in two representations:
1️⃣ A binary index (I used IndexBinaryFlat for exact search and IndexBinaryIVF for approximate search)
2️⃣ An int8 view, i.e., a way to quickly read int8 embeddings from the disk by document ID
➡️ In end, instead of fp32 embeddings, you store:
- a binary index (32 times smaller)
- int8 embeddings (4 times smaller)
Plus, Only the binary index is kept in memory, so the RAM savings are also x32 compared to fp32 search.
For comparison: a regular fp32 retrieval on such a task would require about 180GB of RAM, 180GB of disk for embeddings, and would be 20–25 times slower.
A binary retrieval with int8 rescoring fits into about 6GB of RAM and ~45GB of disk for embeddings.
For example, if you load 4 times more documents through the binary index and then rescoring them in int8, you can recover about 99% of the quality of fp32 search (compared to ~97% for pure binary search)
HFblog
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
