A retrieval-augmented generation (RAG) system has more moving parts than "search plus a language model". Knowing each one makes problems easier to locate.
Offline: Preparing Knowledge
- Sources — documents, wikis, tickets, databases.
- Parsing — turning PDFs, HTML and slides into clean text with structure.
- Chunking — splitting text into retrievable passages.
- Enrichment — titles, headings, dates, permissions and other metadata.
- Indexing — embeddings in a vector index, plus a keyword index.
Online: Answering a Question
- Query understanding — rewriting, expanding or classifying the question.
- Retrieval — finding candidate chunks with semantic and keyword search, filtered by metadata and permissions.
- Re-ranking — ordering candidates precisely.
- Context assembly — choosing chunks and formatting them for the prompt.
- Generation — the model answers from the context, with citations.
- Post-processing — validation, citation checks and guardrails.
Around It
- Evaluation of retrieval and answers on a test set.
- Monitoring of quality, latency and cost.
- Feedback from users to find failures.
- Refresh pipelines keeping the index current.
Where Problems Usually Are
Most poor RAG answers trace back to parsing, chunking or retrieval — not the language model. Measure each stage separately.