Vector Databases
AI revolution makes vector databases very popular. They are now a fundamental block for building GenAI systems. So let's check what they are and how they are different from traditional databases.
Vector database is a database designed to store, manage and index high-dimensional vector data (check Data Vectorization, Embeddings and Word Embeddings for more details). Unlike relational databases this type of database works with unstructured data (embeddings for social media, images, audio, etc.) and can provide search result based on data similarity instead of exact match (e.g., return tomatoes, potatoes and cucumbers on vegetables search request).
Working pipeline consists of 3 steps:
✏️ Indexing. As with other data types, efficiently querying a large set of vectors requires an index. Main algorithms:
- Locality-Sensitive Hashing (LSH) – Uses hashing to group similar vectors.
- Quantization – Compresses data to speed up searches. Compression can loose some initial data but it keeps the info that is is vital for similarity operations.
- Graph-Based Algorithms – Uses nodes to represent vectors. It clusters the nodes and draws lines or edges between similar nodes, creating hierarchical graphs. When a query is launched, the algorithm will navigate the graph hierarchy to find nodes containing the vectors that are most similar to the query vector.
✏️ Search. The system compares the query vector to indexed vectors to find the closest matches:
- Exact Nearest Neighbor - measuring the absolute distance between all points in the vector, it's accurate, slow, requires a lot of computational resources
- Approximate Nearest Neighbor (ANN) - allows to return points whose distance is at most c times the distance from the query to its nearest points. It's cheaper and faster then exact match approach, modern vector databases use ANN algorithms.
✏️ Post (or Pre)-processing. Additional steps may be applied to the search results: metadata filtering, re-ranking based on different similarity measures to improve accuracy.
Popular opensource implementations:
- Opensearch
- Milvus
- Qdrant
- Neo4j
- Pgvector (based on Postgres)
Using LLMs without a vector database can be slow and inefficient because the model must process the full context every time. A vector database optimizes this by storing precomputed embeddings, enabling fast, efficient searches without repeatedly running the entire dataset through the model.
#engineering #aibasics
Post #139
391
- 👍 3