Keeping the Index Fresh: Updates Without Duplicates

Freshness is a key reason to choose RAG over fine-tuning, yet many pipelines can only append. This chapter covers deterministic chunk IDs and upserts, the orphan-chunk trap when documents shrink, content hashing to skip unchanged text, deletions, and full re-indexing when the embedding model changes.

Advanced RAG Masterclass

In Module 1 we chose RAG over fine-tuning partly because knowledge changes: "update the index, not the model." Now hold that promise to account. The 2026 handbook arrives with office hours changed from 9:00–18:00 to 10:00–19:00. You run the ingestion script again. What does the index contain afterwards?

With a naive pipeline, it contains both versions. Random IDs mean every run inserts fresh records, so the old chunk saying 9:00–18:00 and the new one saying 10:00–19:00 sit side by side with nearly identical embeddings. Retrieval returns both, and the model either picks one at random or hedges. The system has become less reliable with each update. Freshness has to be designed in.

The key: deterministic IDs and upsert

Vector databases write with upsert (update or insert): if a record with this ID exists, overwrite it, otherwise create it. Upsert can only overwrite if the same chunk gets the same ID every time it is ingested. So derive the ID from the document itself (document name, section and position), never from a random UUID. For example:

  • employee-handbook_working-hours_chunk-1
  • employee-handbook_working-hours_chunk-2
  • employee-handbook_code-of-conduct_chunk-1

Now re-ingesting the updated handbook produces employee-handbook_working-hours_chunk-1 again, with the new text and a new vector, and the upsert replaces the old record in place. The old hours disappear and the new ones are searchable immediately. No model retraining, no rebuild, just a write.

A random UUID breaks this completely: the database never sees the same ID twice, so every upsert is an insert, and duplicates accumulate silently.

Structure-aware chunking (Module 2) makes good IDs easy, because sections give natural, stable names. Paragraph-level chunks within a section get a counter, and over-long paragraphs are split further and numbered in sequence.

Day 1: handbook chunked into records with deterministic IDs; Day 100: the working-hours section changes; re-ingestion generates the same IDs, so upsert overwrites the changed record in place, while a random-UUID pipeline would insert a duplicate
Deterministic IDs make re-ingestion idempotent: the same chunk always gets the same ID, so an upsert overwrites it. Random IDs turn every update into a duplicate.

The trap: orphan chunks

Deterministic IDs fix changed chunks. They do not fix chunks that no longer exist. Suppose the Working Hours section used to produce three chunks and, after an edit, produces two. Re-ingestion upserts …chunk-1 and …chunk-2, but …chunk-3 from the old version is never touched. It stays in the index as an orphan: stale text that will still be retrieved.

There is a related problem: insert a new paragraph at the top of a section, and every later chunk's number shifts by one, so chunk-2 now holds what used to be chunk-1. Upserts overwrite every chunk with its neighbour's text. Nothing is lost, but every chunk is rewritten and re-embedded.

Production pipelines handle this explicitly:

  1. Replace per document (or per section). When a document changes, delete all records whose ID starts with that document's prefix (or whose doc_id metadata matches), then insert the fresh chunks. It is simple and always correct, and it is fine for documents of normal size.
  2. Keep a manifest. Store the list of chunk IDs produced for each document at the last ingestion. On re-ingest, compute the new list, upsert it, and delete the IDs in the old list that are not in the new one.

Skipping work: content hashing

Re-embedding a 500-page manual because one paragraph changed wastes time and money. Store a hash of each chunk's text in its metadata. On re-ingestion, compute the hashes of the new chunks and compare: unchanged hash means skip, changed or new means embed and upsert, and IDs that have disappeared mean delete. Some teams use the content hash as the ID (so identical text always maps to the same record, and edits naturally create new IDs), combined with document-level deletion of IDs that are no longer present.

Deletions are updates too

When a document is retired, such as a discontinued product or a withdrawn policy, its chunks must be removed: deleted by doc_id filter or ID prefix. Forgetting deletions is the most common source of "the bot quotes a policy we cancelled a year ago." If your source system (SharePoint, Confluence, a CMS) emits change events, subscribe to them, so that creates, updates and deletes flow into the index automatically, rather than relying on periodic full re-scans.

Overwrite or keep history?

Overwriting keeps the index lean and answers always current. Sometimes, though, you must answer "what was the rule in 2025?", for audits, disputes or regulated industries. Then don't overwrite: give each version its own records (include the version in the ID) with year or effective_from and effective_to metadata, and filter for the current version by default (Metadata Filtering covers this). Choose per collection: overwrite for FAQs, version for policies and contracts.

When everything must be re-embedded

Some changes invalidate the whole index:

  • A new embedding model. Vectors from different models are not comparable (Module 2), so every chunk must be re-embedded.
  • A new chunking strategy. Chunk boundaries, and therefore IDs and vectors, all change.

Do these as a blue–green re-index: build a complete new index alongside the live one, run your evaluation set against it, then switch traffic over and keep the old index briefly for rollback. Never re-embed in place while serving traffic, or queries will compare new-model query vectors against a mix of old- and new-model chunks.

Try it yourself
RAG Lab: index freshness →

Re-ingest a changed document under three ID schemes and find the duplicates and orphans.

EasyFreshnessInterview

Why do random UUIDs as chunk IDs cause duplicate, conflicting answers after updates?

MediumFreshness

A section shrinks from three chunks to two. With deterministic IDs and upsert, what goes wrong, and how do you fix it?

HardFreshnessOperations

You are switching embedding models for a live system. How do you do it safely?