Retrieval-augmented generation (RAG) answers questions using your own documents: it retrieves relevant passages and gives them to a language model along with the question.
Why RAG
- Models don't know your internal documents or recent information.
- Grounding answers in retrieved text reduces hallucination and allows citations.
- Updating knowledge means updating documents, not retraining a model.
The Components
- Ingestion: split documents into chunks and store them with metadata.
- Indexing: create embeddings for each chunk and store them in a vector index, often alongside a keyword index.
- Retrieval: embed the user's question and find the most relevant chunks.
- Generation: put those chunks in the prompt with instructions to answer from them and cite sources.
Where It Goes Wrong
| Symptom | Likely cause |
|---|---|
| Says the answer isn't available when it is | Retrieval missed it: chunking, embeddings, or too few results |
| Confident answer not in the documents | Weak grounding instructions or irrelevant chunks |
| Mixes up similar policies | Chunks lack titles, dates or source metadata |
| Wrong after documents change | Stale index |
Improving It
- Chunk along document structure and keep headings with each chunk.
- Combine keyword and semantic search (hybrid search).
- Re-rank retrieved chunks with a stronger model.
- Filter by metadata and user permissions.
Evaluate the Two Halves
Measure retrieval (was the right chunk found?) separately from answer quality (was the answer correct and supported?). Fix retrieval first.