Document collections are full of duplicates: copies in different folders, drafts, attachments forwarded many times, pages repeated across sites.
Why Duplicates Hurt
- The top retrieved results become copies of the same passage, crowding out other useful information.
- Different versions of a document may disagree, producing inconsistent answers.
- They waste storage, embedding cost and context space.
Finding Exact Duplicates
Hash the normalised text of each document or chunk; identical hashes are exact duplicates.
Finding Near-Duplicates
- MinHash and locality-sensitive hashing detect documents sharing most of their content efficiently.
- Embedding similarity finds chunks with near-identical meaning.
Choosing What to Keep
Prefer the canonical source, the most recent version, or the official location. Record the others as aliases so citations can still point users to the right place.
At Retrieval Time
Even with deduplication, similar chunks can appear together. Diversify results with methods such as maximal marginal relevance, which balances relevance with novelty.
Fix It at the Source
Duplicates often reflect content management problems. Work with content owners to archive outdated copies.