TGViewer
TechLead Bits TechLead Bits @techleadbits · 518 subscribers
Post #139 391
Vector Databases

AI revolution makes vector databases very popular. They are now a fundamental block for building GenAI systems. So let's check what they are and how they are different from traditional databases.

Vector database is a database designed to store, manage and index high-dimensional vector data (check Data Vectorization, Embeddings and Word Embeddings for more details). Unlike relational databases this type of database works with unstructured data (embeddings for social media, images, audio, etc.) and can provide search result based on data similarity instead of exact match (e.g., return tomatoes, potatoes and cucumbers on vegetables search request).

Working pipeline consists of 3 steps:

✏️ Indexing. As with other data types, efficiently querying a large set of vectors requires an index. Main algorithms:
- Locality-Sensitive Hashing (LSH) – Uses hashing to group similar vectors.
- Quantization – Compresses data to speed up searches. Compression can loose some initial data but it keeps the info that is is vital for similarity operations.
- Graph-Based Algorithms – Uses nodes to represent vectors. It clusters the nodes and draws lines or edges between similar nodes, creating hierarchical graphs. When a query is launched, the algorithm will navigate the graph hierarchy to find nodes containing the vectors that are most similar to the query vector.

✏️ Search. The system compares the query vector to indexed vectors to find the closest matches:
- Exact Nearest Neighbor - measuring the absolute distance between all points in the vector, it's accurate, slow, requires a lot of computational resources
- Approximate Nearest Neighbor (ANN) - allows to return points whose distance is at most c times the distance from the query to its nearest points. It's cheaper and faster then exact match approach, modern vector databases use ANN algorithms.

✏️ Post (or Pre)-processing. Additional steps may be applied to the search results: metadata filtering, re-ranking based on different similarity measures to improve accuracy.

Popular opensource implementations:
- Opensearch
- Milvus
- Qdrant
- Neo4j
- Pgvector (based on Postgres)

Using LLMs without a vector database can be slow and inefficient because the model must process the full context every time. A vector database optimizes this by storing precomputed embeddings, enabling fast, efficient searches without repeatedly running the entire dataset through the model.

#engineering #aibasics
  • 👍 3
More from @techleadbits
  1. Oct 7, 2026AI & Repository Strategy For many years, there has been an ongoing debate between monorepo…
  2. Oct 1, 2026Tracer Bullets Continuing the topic from the previous post, let's talk in more detail abou…
  3. Sep 28, 2026Why Software Factories Fail "Read the Code!" is one of the key ideas from Dex Horthy's tal…
  4. Sep 21, 2026Illustrations from The Culture Map showing how different cultures compare on the scales. #…
  5. Sep 21, 2026The Culture Map Have you ever worked in international distributed teams? Or collaborated w…
  6. Sep 10, 2026Loop Engineering from First Principles Continuing the topic of Loop Engineering, I'd like…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →