A RAG system can only retrieve what it parsed correctly. Poor parsing produces garbled chunks that no embedding model can rescue.
Common Problems
- PDFs: multi-column layouts read in the wrong order, headers and footers repeated on every page, hyphenated line breaks, tables flattened into nonsense.
- Scans: need OCR, with errors in unusual fonts or low-quality images.
- Slides: text spread across boxes with little structure.
- HTML: navigation, cookie banners and footers mixed in with content.
- Spreadsheets: meaning lives in rows, columns and headers together.
Preserve Structure
Keep headings, lists and tables. Structure helps chunking along natural boundaries, and headings give each chunk context.
Tables
Convert tables to Markdown or to one line per row with column names ("Plan: Pro, Price: $20, Seats: 5"). Very large tables may be better queried from a database than retrieved as text.
Images and Diagrams
Diagrams often carry key information. A multimodal model can describe them in text for indexing, or you can index captions and surrounding text.
Clean Boilerplate
Strip repeated headers, footers, page numbers, legal disclaimers and navigation.
Check a Sample
Read the parsed output of a random sample of documents before indexing everything. Parsing errors are easy to spot by eye and expensive to discover later.