TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 582 subscribers
Post #1967 201
Binary search + rescoring in int8.

The simple strategy to search through 40 million texts in ~200 ms using only a CPU server, 8GB of RAM, and 45GB of disk space.

If you want to try it out immediately, there's a demo of 40 million texts from Wikipedia. No login or other hassles required.

🟢 The Inference Strategy:

☞ Embed the query with a dense model into a regular fp32 vector

☞ Quantize the fp32 embedding into a binary format, which is 32 times smaller

☞ Retrieve, for example, 40 documents (about 20 times faster than an fp32 index) using an approximate or exact binary index

☞ Load the int8 embeddings for these top-40 documents from the disk

☞ Rescoring: fp32 embedding of the query × 40 int8 embeddings

☞ Sort these 40 documents by the new score and take the top-10

☞ Load the titles and texts of the top-10 documents

The documents are embedded once, and then these embeddings are used in two representations:

1️⃣ A binary index (I used IndexBinaryFlat for exact search and IndexBinaryIVF for approximate search)

2️⃣ An int8 view, i.e., a way to quickly read int8 embeddings from the disk by document ID

➡️ In end, instead of fp32 embeddings, you store:

- a binary index (32 times smaller)
- int8 embeddings (4 times smaller)

Plus, Only the binary index is kept in memory, so the RAM savings are also x32 compared to fp32 search.
For comparison: a regular fp32 retrieval on such a task would require about 180GB of RAM, 180GB of disk for embeddings, and would be 20–25 times slower.

A binary retrieval with int8 rescoring fits into about 6GB of RAM and ~45GB of disk for embeddings.

For example, if you load 4 times more documents through the binary index and then rescoring them in int8, you can recover about 99% of the quality of fp32 search (compared to ~97% for pure binary search)


HFblog
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →