Vector retrieval finds passages similar to a question. It struggles with questions about relationships across many documents or themes of a whole collection.
What a Knowledge Graph Adds
A knowledge graph stores entities (people, products, organisations) and relationships between them (works for, depends on, supplies). It can answer "which suppliers of component X also supply competitors?" by traversing connections.
GraphRAG Approaches
- Entity-centric retrieval: extract entities from the question, find them in the graph, and retrieve connected entities and source passages.
- Community summaries: cluster the graph into communities of related entities and pre-generate summaries, enabling answers to broad questions like "what are the main themes in these reports?"
- Hybrid: combine graph traversal with vector search.
Building the Graph
Language models can extract entities and relationships from text, but extraction errors and duplicate entities are common. Entity resolution — deciding that "IBM" and "International Business Machines" are one node — is essential.
Costs
Building and maintaining a graph adds significant processing and complexity, and updating it as documents change takes care.
When It's Worth It
Relationship-heavy domains (supply chains, compliance, investigations, research literature) and questions about whole corpora. For factual lookups in documentation, standard hybrid RAG is usually enough.
Evaluate the Difference
Test with relationship and summary questions specifically, comparing against standard RAG.