RAG means Retrieval-Augmented Generation. You split your documents into small chunks, turn each chunk into a vector with an embedding model, and store the vectors in a database like pgvector or Qdrant. When a question comes, you embed it, fetch the closest chunks, the top 3 to 5, and put them in the prompt. The model answers using only that text. With LlamaIndex or LangChain, a basic version is about 30 lines of Python.
Most weak RAG systems fail at retrieval, not at writing. If the right chunk is never found, even the best model can only guess. Fix retrieval first.
Three things decide quality.
Chunking. Split on paragraphs or headings, not by raw character count. Keep chunks near 300 to 500 tokens, with about 15 percent overlap so sentences stay whole.
Embedding model. A weak one retrieves the wrong chunks. For English, Gemini Embedding 001 or OpenAI text-embedding-3-large are strong. For open, multilingual use, BGE-M3 and Qwen3-Embedding are 2026 defaults.
Retrieval testing. Write 30 real questions, mark the chunk that should answer each, and measure how often it reaches the top. That number is recall@k, and Ragas can track it for you. This habit finds most bugs.
Two upgrades give a big jump. Hybrid search runs keyword search (BM25) and vector search, then merges them with Reciprocal Rank Fusion (RRF), so exact names, codes, and IDs are not missed. A reranker like Cohere Rerank 3.5 or the open BGE-reranker-v2-m3 then keeps the best 5 of the top 20.
One more rule. Tell the model to answer only from the given text and to say it does not know when the answer is missing. Keep a source link on every chunk, so people can check it.
RAG vs fine-tuning, in one line: RAG adds fresh facts you can update any day; fine-tuning shapes style and behavior. For your own documents, start with RAG.