A RAG request involves several steps, each adding latency and cost. Knowing where they go lets you optimise sensibly.
Where Time Goes
- Query rewriting (a model call).
- Embedding the query.
- Vector and keyword search.
- Re-ranking.
- Generating the answer — usually the largest share.
Reducing Latency
- Run keyword and vector searches in parallel.
- Use a small, fast model for query rewriting, or skip it when the question is already standalone.
- Re-rank a moderate number of candidates.
- Send fewer, more relevant chunks to the model.
- Stream the answer to the user.
- Cache embeddings of common queries and answers to frequent questions.
Reducing Cost
- Use the smallest generation model that meets quality targets.
- Trim context: fewer chunks, less boilerplate.
- Use prompt caching for stable instructions.
- Batch embedding during ingestion.
- Route simple questions to cheaper paths.
Measure Per Stage
Log timing and token counts for each stage. Optimise the biggest contributors first.
Protect Quality
Re-run your evaluation after every optimisation. Faster, cheaper answers that are wrong cost more in the end.