RAG Returns Stale Chunks After a Document Update: Versioned Ingestion
AI Tooling · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your vector store keeps serving text that was edited or deleted weeks ago. Here's the ingestion architecture that makes updates and deletes actually stick.
The Forces at Play
When a RAG system returns stale chunks after a document update, it is almost never the model's fault — it is the ingestion pipeline quietly leaving two generations of the same document in the index. Someone edits a policy page, the nightly job runs green, and the assistant still cites the sentence that was deleted, because retrieval is happily ranking both versions against the query.
Three forces pull against each other here:
- Vector stores are key-value upserts, not transactions. There is no
UPDATE ... WHERE doc = xsemantics you get for free. A write touches exactly the IDs you name. - Chunking is non-deterministic across code changes. Chunk count is a function of content and of your splitter config. Any edit to either changes the ID space.
- Retrieval has no notion of freshness. Cosine similarity does not care that a chunk is orphaned. Worse, the old paragraph is often the better lexical match, because the new one was rewritten in different words.
The naive design — "embed chunks, upsert by {doc_id}-{index}" — satisfies none of this. It is correct only in the case where every document's chunk count never shrinks and your splitter never changes. Neither holds.
The Shape
The fix is to stop treating chunks as individually mutable rows and start treating a document version as an immutable set, with a small pointer table deciding which set is live.
flowchart TB
SRC[Source of truth<br/>CMS / S3 / Confluence] -->|content + updated_at| DET{Change detector<br/>content hash}
DET -->|hash unchanged| SKIP[No-op]
DET -->|hash changed| PLAN[Build version id<br/>v = sha256 content+splitter_cfg]
PLAN --> CHUNK[Chunker]
CHUNK --> EMB[Embedder<br/>batched]
EMB --> WRITE[(Vector store<br/>chunks tagged doc_id + doc_version)]
WRITE --> PTR[(Active version table<br/>doc_id -> doc_version)]
PTR -.->|flip after all chunks written| WRITE
PTR --> SWEEP[Sweeper<br/>delete doc_version != active]
SWEEP --> WRITE
Q[User query] --> RET[Retriever]
PTR -->|filter: active versions| RET
RET -->|top-k within active set| WRITE
RET --> LLM[Answer + citations]
The critical arrows are the dotted one and the filter. Chunks become visible only after the pointer flip, and the retriever filters on the active version, so an orphaned set is unreachable even before the sweeper runs. Deletion becomes a garbage-collection concern rather than a correctness concern.
How Data Flows Through It
Legal edits policy/refunds, changing "30 days" to "14 days".
- Detect. The job pulls the doc and computes
sha256(normalized_body + splitter_config_version + embed_model_id). It differs from the stored hash, so this is a real change — not a whitespace-only CMS republish. (Including the splitter and model ID means a config change re-ingests everything automatically.) - Plan. That hash becomes
doc_version = "a91f...". Nothing is deleted yet. - Chunk and embed. The doc yields 8 chunks; the previous version yielded 12. Each chunk carries
doc_id,doc_version,ordinal, and a source offset for citation. - Write. All 8 chunks are inserted with IDs
a91f...-0..7. They coexist with the 12 old chunks. Retrieval still serves the old version — correct, because the new set may be half-written. - Flip. One row update:
active_version['policy/refunds'] = 'a91f...'. This is the commit point. Retrieval now filters to the new set; all 12 old chunks vanish from results instantly and atomically. - Sweep. An async job deletes
doc_version != active. If it crashes, nothing is user-visible — the pointer already guarantees correctness. You just pay storage. - Query. "How long do I have to request a refund?" hits the retriever, which injects
doc_version IN (active set)as a metadata filter and returns top-k from live chunks only.
The message contract between the ingest worker and the store is worth pinning down:
{
"id": "a91f3c...-0",
"doc_id": "policy/refunds",
"doc_version": "a91f3c...",
"ordinal": 0,
"embed_model": "text-embedding-3-large",
"splitter_cfg": "recursive-800-100-v3",
"source_span": [0, 812],
"text": "Customers may request a refund within 14 days..."
}
Note id is derived from the version hash, not from position. Two versions can never collide, so a partial write can never corrupt the live set.
What Each Piece Owns
Change detector owns deciding whether work is needed. It does not own chunking policy — it only needs a stable hash that includes every input that could change output. Leaving the splitter config out of that hash is the classic bug: you tune chunk size, redeploy, and nothing re-ingests.
Chunker owns text boundaries and source offsets for citations. It does not own IDs or deduplication; it emits an ordered list and nothing else.
Embedder owns batching, retry, and rate-limit backoff against the model API. It does not own partial-failure semantics for the document — it either returns all vectors for a version or raises, so the pointer never flips on a half-embedded set.
Vector store owns similarity search and metadata filtering. It explicitly does not own freshness or the notion of "current". Every stale-chunk bug comes from a team assuming it does.
Active version table owns the commit point. It is the only transactional component, and it is deliberately tiny — one row per document, so it can live in your existing relational DB even when the vectors don't.
Sweeper owns storage cost. It owns nothing about correctness, which is exactly why it's allowed to be a best-effort cron.
Where It Breaks Down
The metadata filter becomes the bottleneck first. Filtering by a high-cardinality doc_version across millions of vectors can force pre-filtering that degrades ANN recall, or post-filtering that returns fewer than k results. In HNSW-backed stores, an aggressive filter that excludes most of the graph makes search fall back to near-brute-force. Mitigation: keep the sweeper aggressive enough that stale vectors are a small fraction, and filter on a cheap boolean is_active maintained by the sweeper rather than an IN-list of thousands of hashes.
Partial failure hurts on multi-doc atomicity. Per-document commit is easy; "this 200-page handbook must update as a unit" is not. If half the handbook flips and half doesn't, the LLM will happily synthesize the two into a contradiction. If you need that, add a collection-level version and flip once.
Re-embedding the corpus is the real operational burden. Changing embedding models means every document gets a new version simultaneously — double storage, a large API bill, and a window where both sets exist. Plan for a store that can hold 2× your corpus, and roll the flip out per-collection so you can A/B retrieval quality before committing.
Deletes at source are silently missed. If a document is removed from the CMS, no change event fires from a pull-based detector that only iterates existing docs. You need a reconciliation pass: list all doc_ids in the store, diff against the source listing, and tombstone the difference. Do this weekly; stale deleted documents are the ones that cause compliance incidents.
When This Is Overkill
If your corpus is under a few thousand documents and rebuild takes minutes, just rebuild the whole index into a new collection and swap an alias. One name, two collections, atomic pointer flip at the collection level — the same idea with far less machinery, and zero orphan risk because you never mutate anything.
# build into refunds_kb_2024_06_11, verify, then:
vector-cli alias set refunds_kb -> refunds_kb_2024_06_11
You've outgrown full rebuilds when any of these appear:
- Rebuild time exceeds your acceptable staleness window (docs change hourly, rebuild takes 90 minutes).
- Re-embedding the full corpus on every run costs more than the feature is worth.
- You have per-tenant or per-ACL documents that must update independently without touching each other's data.
Until then, per-document versioning is machinery you get to not operate. The one thing you should adopt immediately regardless of scale: never use positional chunk IDs. That single choice is the difference between "delete works" and "we've been citing a deleted policy for three weeks."
Key takeaway: Treat a document's chunks as one immutable versioned set: write the new version, flip the active pointer, then delete the old version — never upsert chunk-by-chunk by positional ID.
Real-world challenge
Support escalates: the assistant keeps quoting a refund window of '30 days' even though legal updated the policy page to 14 days three weeks ago. The ingestion job runs nightly, logs show zero errors, and querying the vector store for the policy doc does return a chunk containing '14 days'. Yet answers still cite 30. What's wrong and how do you fix it?
Diagnose. Don't trust the ingestion logs — count what's actually stored:
-- pgvector example
SELECT doc_version, count(*), min(ingested_at)
FROM chunks WHERE doc_id = 'policy/refunds'
GROUP BY doc_version;
If you see two doc_version values (or chunks with two different ingested_at days), both the old and new chunk sets are live. Retrieval returns top-k across all of them, and the old '30 days' chunk is often more lexically similar to the question than the rewritten paragraph.
Root cause. The job upserts {doc_id}-{index} IDs. The rewrite shortened the document, so the trailing chunks from the previous version were never overwritten and never deleted.
Fix (two parts).
- Make retrieval version-aware immediately — filter on the active version so stale rows can't be returned even before cleanup:
results = store.query(vec, k=8, filter={"doc_version": active[doc_id]})
- Change ingestion to write-then-sweep: insert the new version's chunks, atomically update the
active_versionrow for that doc, thenDELETE FROM chunks WHERE doc_id=$1 AND doc_version <> $2.
Prevent regression. Add a post-ingest assertion that count(distinct doc_version) = 1 per doc and alert on any doc whose oldest chunk predates its source updated_at.