An algorithm from the 90s, without training, without embeddings and fine-tuning, still forms the basis of Elasticsearch, OpenSearch and most modern search engines. It's called BM25 and worth understanding why it's still going strong.
Let's say you're searching for "transformer attention mechanism" in a library of ML articles.
🟢 BM25 assesses document relevance based on three core ideas:
1️⃣ Rare words matter more than common ones
Every article contains "the" and "is", so such words carry little weight.
But "transformer" is specific and informative, so BM25 gives it significantly more weight. This is reflected in the formula through IDF(qᵢ).
2️⃣ Repetitions help, but with diminishing returns
If "attention" appears 10 times in an article, it's a strong signal of relevance. But increasing it from 10 to 100 mentions barely changes the final score.
BM25 uses saturation, controlled by f(qᵢ, D) and the parameter k₁, to prevent keyword stuffing from artificially boosting rankings.
3️⃣ Document length is normalized
A50-page article will naturally have more keyword occurrences than a 5-page one.
BM25 accounts for this through |D|/avgdl, controlled by the parameter b, to ensure long documents don't dominate the results just because they have more text.
Three ideas.
Zero neural networks.
Zero datasets.
Just meticulous mathematics that's stood the test of time.
➡️ WHAT MANY OVERLOOK?
BM25 excels at exact keyword matches, while embeddings can struggle with this.
When a user searches for "error code 5012", vector search might return semantically similar error codes. BM25 almost always prioritizes the exact match at the top.
That's why hybrid search has become the default in top-tier RAG systems.
The combination of BM25 + vector search delivers both semantics and exact keyword matches in a single pipeline.
So before throwing a GPU at any search task, consider: maybe BM25 already solves it. And if not, it'll almost certainly improve your semantic search significantly in tandem.
This hybrid search stack I mentioned above is actually already implemented in the open-source layer of the contextual retriever for agents.
GitHub
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
