First-stage retrieval casts a wide net; a re-ranker picks the best catch.
Why Re-Rank
Embedding search compares separately computed vectors, which is fast but approximate. A re-ranker examines the question and each candidate together and scores relevance much more precisely.
Types of Re-Rankers
- Cross-encoders: transformer models that take the question and passage as a pair. Accurate and reasonably fast for tens of candidates.
- LLM re-rankers: a language model scores or orders passages. Flexible and strong, but slower and more expensive.
- Hosted re-ranking APIs offered by several providers.
Choosing the Candidate Count
Retrieve, say, 50 candidates and re-rank to keep the best 5. More candidates improve recall but increase re-ranking time. Tune using your test set.
Measuring the Benefit
Compare recall@k and final answer quality with and without re-ranking. Re-ranking often gives one of the largest single improvements in RAG quality.
Domain Fit
General re-rankers may struggle with specialised language. Test on your content; fine-tuning with domain question–passage pairs can help.
Latency Budget
Re-ranking adds a step to every request. Run it on a GPU or use a smaller model if latency matters, and measure end-to-end.