Retrieval works on passages, not whole documents. How you split documents into chunks often decides whether a RAG system works.
Why Chunk
An embedding of a 40-page manual is a blurry average of everything in it. Smaller chunks give precise matches and let you pass only relevant text to the language model.
Choosing a Size
- Too small (a sentence): chunks lose the context needed to understand them.
- Too large (many pages): embeddings blur and prompts fill with irrelevant text.
- A few hundred words is a common starting point, and must fit the embedding model's input limit. Test sizes against real questions.
Split Along Structure
Split by headings, sections and paragraphs rather than a fixed character count. Keep tables and lists intact where possible.
Overlap
Overlapping neighbouring chunks slightly helps when an answer straddles a boundary.
Add Context and Metadata
- Prefix each chunk with its document title and section heading.
- Store source, URL, date, version and access permissions as metadata for filtering and citations.
Special Content
- Tables: convert to text rows or keep as small, self-contained chunks.
- Code: split by function or class.
- Scanned PDFs: need OCR first; check quality.
Keep It Fresh
Re-index changed documents, delete chunks from removed ones, and record which version was indexed.
Test It
Write realistic questions with known answers and check whether the right chunk appears in the top results before involving a language model.