Skip to content

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

Editorial team 1 min read

A retrieval-augmented generation (RAG) system has more moving parts than "search plus a language model". Knowing each one makes problems easier to locate.

Offline: Preparing Knowledge

  1. Sources — documents, wikis, tickets, databases.
  2. Parsing — turning PDFs, HTML and slides into clean text with structure.
  3. Chunking — splitting text into retrievable passages.
  4. Enrichment — titles, headings, dates, permissions and other metadata.
  5. Indexing — embeddings in a vector index, plus a keyword index.

Online: Answering a Question

  1. Query understanding — rewriting, expanding or classifying the question.
  2. Retrieval — finding candidate chunks with semantic and keyword search, filtered by metadata and permissions.
  3. Re-ranking — ordering candidates precisely.
  4. Context assembly — choosing chunks and formatting them for the prompt.
  5. Generation — the model answers from the context, with citations.
  6. Post-processing — validation, citation checks and guardrails.

Around It

  • Evaluation of retrieval and answers on a test set.
  • Monitoring of quality, latency and cost.
  • Feedback from users to find failures.
  • Refresh pipelines keeping the index current.

Where Problems Usually Are

Most poor RAG answers trace back to parsing, chunking or retrieval — not the language model. Measure each stage separately.

More in RAG

All RAG guides →
RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025

RAG Guide · 1 min

Query Rewriting and Expansion

Improving retrieval by rewriting user questions: expansion, decomposition, hypothetical answers and conversational context.

RAG 1 min read 2 Dec 2025