Skip to content

Re-Ranking Search Results

How a second-stage re-ranker improves retrieval quality for search and RAG, and the trade-offs in speed and cost.

Editorial team 2 min read

Fast retrieval methods — keyword search and vector search — are good at finding plausible candidates but imperfect at ordering them. A re-ranker fixes the order.

Two-Stage Retrieval

  1. Retrieve a broad set of candidates quickly, say the top 50–100.
  2. Re-rank those candidates with a slower, more accurate model, and keep the best few.

How Re-Rankers Work

Embedding search compares a query vector with document vectors computed separately. A cross-encoder re-ranker reads the query and each candidate together, so it can judge relevance much more precisely. Language models can also be used as re-rankers by asking them to score or order passages.

Benefits

  • Better top results, which matters most when only a few passages go into an LLM prompt.
  • Recovers relevant documents that the first stage ranked too low.
  • Lets you retrieve more broadly without flooding the prompt.

Costs

Re-ranking adds latency and compute proportional to the number of candidates. Keep the candidate set moderate and measure the end-to-end effect.

Where It Helps Most

Question answering over large or varied document collections, customer support search, and RAG systems where irrelevant context leads to poor answers.

Measure It

Compare metrics such as recall@k and the rate of correct final answers with and without re-ranking on a realistic test set.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026