ChatGPT SearchSep 27, 2026
The key is to make document changes event-driven, rather than periodically re-embedding the entire corpus.
Recommended architecture
Source docs
│
├── change detected (webhook / CDC / Git commit / file event)
▼
Change queue
│
▼
Ingestion worker
├── fetch changed document
├── compare content hash/version
├── chunk
├── embed only changed chunks
└── upsert/delete vectors
▼
Vector DB
│
▼
RAG retrievalThe important design detail is to give every document and chunk a stable ID, plus metadata such as:
document_id
chunk_id
content_hash
document_version
updated_at
source_uriWhen a source changes:
- Calculate its new content hash.
- If unchanged, do nothing.
- If changed, re-chunk that document.
- Re-embed only new/modified chunks.
- Delete vectors belonging to chunks that disappeared.
- Upsert the new vectors.
- Record the new document version.
This prevents expensive full-corpus reindexing.
Tools I'd consider
If you're already using a vector database, its native update/upsert mechanism is usually the right place to start. For example, Weaviate automatically re-vectorizes an object when an existing vectorized property is updated, and its asynchronous indexing can put updates, deletes, and inserts through a persistent queue.