TGViewer
How AI Helps How AI Helps @howaihelps · 794 subscribers
Post #269 380
RAG lets any LLM answer from your own documents, and three small fixes decide whether the answers are good or useless

RAG means Retrieval-Augmented Generation. You split your documents into small chunks, turn each chunk into a vector with an embedding model, and store the vectors in a database like pgvector or Qdrant. When a question comes, you embed it, fetch the closest chunks, the top 3 to 5, and put them in the prompt. The model answers using only that text. With LlamaIndex or LangChain, a basic version is about 30 lines of Python.

Most weak RAG systems fail at retrieval, not at writing. If the right chunk is never found, even the best model can only guess. Fix retrieval first.

Three things decide quality.

Chunking. Split on paragraphs or headings, not by raw character count. Keep chunks near 300 to 500 tokens, with about 15 percent overlap so sentences stay whole.

Embedding model. A weak one retrieves the wrong chunks. For English, Gemini Embedding 001 or OpenAI text-embedding-3-large are strong. For open, multilingual use, BGE-M3 and Qwen3-Embedding are 2026 defaults.

Retrieval testing. Write 30 real questions, mark the chunk that should answer each, and measure how often it reaches the top. That number is recall@k, and Ragas can track it for you. This habit finds most bugs.

Two upgrades give a big jump. Hybrid search runs keyword search (BM25) and vector search, then merges them with Reciprocal Rank Fusion (RRF), so exact names, codes, and IDs are not missed. A reranker like Cohere Rerank 3.5 or the open BGE-reranker-v2-m3 then keeps the best 5 of the top 20.

One more rule. Tell the model to answer only from the given text and to say it does not know when the answer is missing. Keep a source link on every chunk, so people can check it.

RAG vs fine-tuning, in one line: RAG adds fresh facts you can update any day; fine-tuning shapes style and behavior. For your own documents, start with RAG.
  • 👍 2
  • ❤ 1
More from @howaihelps
  1. Oct 6, 2026Mistral launches Large 4 in API preview, leading a test of finding and fixing security bug…
  2. Oct 6, 2026Codex can keep working without another “keep going” An AI assistant investigates a bug, fi…
  3. Oct 6, 2026Narrow down a bug’s cause with a Claude Code team When a bug has several plausible causes,…
  4. Oct 6, 2026HackerRank's Chakra AI interviewer leaves beta HackerRank says Chakra has interviewed over…
  5. Oct 5, 2026Turn a pixel image into an editable vector Recraft’s AI vectorizer converts an existing ra…
  6. Oct 5, 2026AI may reach the person everyone has stopped arguing with At a family dinner, someone anno…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →