The repository contains the LIMIT dataset, created to test embedding models on theoretical principles. The study shows that even modern models cannot retrieve certain documents, highlighting the limitations of the current approach using single-vector embeddings.
🚀Key points:
- Dataset for testing embedding models.
- Includes 50k documents and 1000 queries.
- Highlights theoretical limitations of information retrieval.
- Code for data generation and experiments is available in the repository.
GitHub
#python
🤖 Data Science, ML & Big Data with @DataXplore
