Approximate nearest-neighbour (ANN) indexes make vector search fast by trading a little accuracy. Their parameters control that trade.
HNSW
A graph-based index widely used for its speed and recall.
- M: connections per node. Higher improves recall and memory use.
- ef_construction: search breadth when building. Higher builds a better graph, slower.
- ef_search: search breadth at query time. Higher improves recall, slower queries. Often the main knob to tune.
IVF
Clusters vectors and searches only the nearest clusters.
- Number of clusters (lists): more clusters, faster but may miss results.
- nprobe: clusters searched per query. Higher improves recall, slower.
Quantisation
Compressing vectors (product quantisation, scalar quantisation) reduces memory substantially with some accuracy loss. Useful for very large collections.
How to Tune
- Compute exact nearest neighbours for a sample of queries as ground truth.
- Measure ANN recall against it at different settings.
- Pick settings that meet your recall target at acceptable latency.
Filtering Interacts
Heavy metadata filtering can reduce effective recall in some ANN implementations. Test with realistic filters.
Don't Over-Engineer
For collections under a few hundred thousand vectors, defaults or even exact search are often fine.