Skip to content

Re-Ranking Models for RAG

Cross-encoders and LLM re-rankers: how they sharpen retrieval, and how to decide how many candidates to re-rank.

Editorial team 1 min read

First-stage retrieval casts a wide net; a re-ranker picks the best catch.

Why Re-Rank

Embedding search compares separately computed vectors, which is fast but approximate. A re-ranker examines the question and each candidate together and scores relevance much more precisely.

Types of Re-Rankers

  • Cross-encoders: transformer models that take the question and passage as a pair. Accurate and reasonably fast for tens of candidates.
  • LLM re-rankers: a language model scores or orders passages. Flexible and strong, but slower and more expensive.
  • Hosted re-ranking APIs offered by several providers.

Choosing the Candidate Count

Retrieve, say, 50 candidates and re-rank to keep the best 5. More candidates improve recall but increase re-ranking time. Tune using your test set.

Measuring the Benefit

Compare recall@k and final answer quality with and without re-ranking. Re-ranking often gives one of the largest single improvements in RAG quality.

Domain Fit

General re-rankers may struggle with specialised language. Test on your content; fine-tuning with domain question–passage pairs can help.

Latency Budget

Re-ranking adds a step to every request. Run it on a GPU or use a smaller model if latency matters, and measure end-to-end.

More in RAG

All RAG guides →
RAG Guide · 1 min

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

RAG 1 min read 6 Dec 2025

RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025