* does not require a vector DB
* does not create embeddings
* does not chunk documents
* does not perform similarity search
And it demonstrated 98.7% accuracy on a financial benchmark (SOTA).
🟢 What are the key problem of classical RAG that this approach solves?:
Traditional RAG chunks documents, converts them into vectors, and retrieves fragments based on semantic similarity.
But similarity ≠ relevance.
When you ask: "What were the debt trends in 2023?", vector search will return pieces that are semantically similar to the query.
But the real answer might be hidden somewhere in the Appendix, referenced by a link on another page, in a section that does not overlap semantically with your question at all.
Classical RAG most likely just won't find it.
PageIndex addresses this.
Instead of chunking and embeddings, PageIndex builds a hierarchical tree of the document structure, essentially a smart "table of contents."
Then the model reasons through this tree.
That is, it does not ask: "Which text is most similar to my query?"
It asks: "Based on the document structure, where would a human expert look for the answer?"
This is a fundamentally different approach, which has:
* no arbitrary chunking that breaks context
* no need to maintain and operate a vector DB
* retrieval is traceable: you can see why a particular section was chosen
* it can properly follow internal document links ("see Table 5.3"), just like a person does
But the deeper issue is this.
Vector search treats each query as independent.
Documents have structure and logic: sections refer to each other, context accumulates across pages.
PageIndex respects this structure instead of flattening everything into embeddings.
Importantly: this approach is not always sensible, because classical vector search is still fast, simple, and works great in many cases.
But for professional documents requiring domain expertise and multi-step reasoning, the tree-based, reasoning-first approach truly shines.
For example, PageIndex showed 98.7% accuracy on FinanceBench and significantly outperformed traditional vector-based RAG systems in analyzing complex financial documents.
Everything is fully open source, you can check the implementation on GitHub and try it yourself.
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore