Skip to content

Query Rewriting and Expansion

Improving retrieval by rewriting user questions: expansion, decomposition, hypothetical answers and conversational context.

Editorial team 1 min read

Users rarely phrase questions the way documents are written. Transforming the query before retrieval can improve recall considerably.

Conversational Context

In a chat, "What about for contractors?" means nothing on its own. Rewrite follow-up questions into standalone queries using the conversation history before searching.

Expansion

Add synonyms, acronyms and related terms: "PTO" → "paid time off, annual leave, vacation". A language model can generate variations, or you can maintain a domain glossary.

Multi-Query Retrieval

Generate several phrasings of the question, retrieve for each, and merge results. This catches documents that match one phrasing but not another.

Decomposition

Complex questions — "Compare the refund policies for annual and monthly plans" — can be split into sub-questions retrieved separately.

Hypothetical Document Embeddings (HyDE)

Ask a model to write a hypothetical answer, then embed that answer and search with it. The hypothetical answer often resembles real passages more closely than the question does. Watch that invented details don't mislead retrieval.

Costs

Each rewriting step adds latency and tokens. Use lightweight models for rewriting and apply techniques only where they improve measured recall.

Keep the Original

Always pass the user's original question to the generation step, so the answer addresses what was actually asked.

More in RAG

All RAG guides →
RAG Guide · 1 min

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

RAG 1 min read 6 Dec 2025

RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025