Much knowledge lives in charts, diagrams, screenshots, scanned forms and slides. Text-only RAG misses it.
Approaches
- Convert to text: use a multimodal model or OCR to describe images, extract chart data and transcribe tables, then index the text. Simple and compatible with existing pipelines.
- Multimodal embeddings: embed images and text into a shared space so a text question can retrieve an image directly.
- Page-level retrieval: embed whole page images (for slides and PDFs) and give the retrieved pages to a multimodal model to answer from.
Answering
Multimodal language models can read retrieved images, charts and pages alongside text, answering questions such as "what was the Q3 value in the revenue chart?"
Challenges
- Extracting accurate numbers from charts.
- Small text in images.
- Storage and cost for image embeddings and multimodal generation.
- Evaluating answers that depend on visual detail.
Practical Tips
- Keep captions and surrounding text with images.
- Store page references so answers can link to the exact page.
- Verify numbers extracted from charts when they matter.
Start With the Content That Matters
Identify where important information is visual-only, and add multimodal handling there rather than everywhere.