Skip to content

RAG Latency and Cost Optimisation

Where time and money go in a RAG pipeline, and practical ways to make it faster and cheaper.

Editorial team 1 min read

A RAG request involves several steps, each adding latency and cost. Knowing where they go lets you optimise sensibly.

Where Time Goes

  • Query rewriting (a model call).
  • Embedding the query.
  • Vector and keyword search.
  • Re-ranking.
  • Generating the answer — usually the largest share.

Reducing Latency

  • Run keyword and vector searches in parallel.
  • Use a small, fast model for query rewriting, or skip it when the question is already standalone.
  • Re-rank a moderate number of candidates.
  • Send fewer, more relevant chunks to the model.
  • Stream the answer to the user.
  • Cache embeddings of common queries and answers to frequent questions.

Reducing Cost

  • Use the smallest generation model that meets quality targets.
  • Trim context: fewer chunks, less boilerplate.
  • Use prompt caching for stable instructions.
  • Batch embedding during ingestion.
  • Route simple questions to cheaper paths.

Measure Per Stage

Log timing and token counts for each stage. Optimise the biggest contributors first.

Protect Quality

Re-run your evaluation after every optimisation. Faster, cheaper answers that are wrong cost more in the end.

More in RAG

All RAG guides →
RAG Guide · 1 min

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

RAG 1 min read 6 Dec 2025

RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025