A customer chatbot needs more than a language model. It needs the right store information, a way to update that information, and a reliable answer to a basic question: is the new content actually ready to search?
I’m Jimmy, the founder and developer of ProxyAI. While building its retrieval-augmented generation (RAG) pipeline, I found that coordinating document ingestion deserved as much attention as retrieval itself.
Here is how I structured that part of the system.
Start with a clear boundary
RAG retrieves relevant information and supplies it as context for an answer. For ecommerce, that can include shipping policies, return instructions, and other store content.
I treat this knowledge as a document collection that can change. A merchant editing a policy should replace the existing document, rather than leave several conflicting versions available to search.
This article focuses on that ingestion layer. Live inventory, customer-specific orders, and permissions are separate concerns. A document index should not be treated as the authority for rapidly changing or private transaction data.
Upload once, then hand the job over
The ingestion flow looks like this:
Document upload to R2
|
Manifest creation event
|
Cloudflare Queue
|
Worker claims the job
|
Replace document in AI Search
|
Durable Object tracks indexing
|
Settle job and notify dashboard
The document goes to object storage. A manifest marks the point when the upload is ready for background processing. Its creation triggers the queue consumer.
That handover matters. Once the manifest is written, the queue owns the job. The merchant does not need to keep the upload page open while indexing finishes.
The worker claims the job through the application backend. The backend’s job record supplies the authoritative document identity and upload details. A browser-written manifest is a trigger, not permission to choose arbitrary content or a destination.
Give documents stable identities
In ProxyAI, the document name is an identity, not just a display filename. An edited document can use the same identity, and a synced WordPress page can use its permalink.
Before uploading the replacement, the consumer deletes the stored version. This prevents an ordinary content update from accumulating obsolete copies.
There is a tradeoff: delete followed by upload is not an atomic swap. If the upload fails, the replacement job needs to be retried. Stable identity helps make that retry understandable, but it does not remove every failure mode.
Let the search service own chunking
I upload the document whole to Cloudflare AI Search and let it handle chunking and embedding. That reduces the number of intermediate objects and queue messages my application needs to coordinate.
AI Search also supports hybrid retrieval, combining semantic and keyword matching. That is useful when questions mix natural language with exact product terms. See Cloudflare’s AI Search documentation for the current capabilities.
The tradeoff is that chunking becomes part of the search service’s configuration. Changing that configuration can change retrieval results, even when the source text stays the same.
Separate uploading from waiting
The queue worker uploads the document, then hands tracking to an IngestJob Durable Object. It does not stay in a long loop waiting for indexing.
The object stores the job’s state and uses alarms to check progress. It renews the job lease while work is still underway, then reports completion or failure to the backend.
This is still polling the indexing service. The improvement is where the waiting happens and how its state survives interruptions.
Reconnection needs a snapshot
The dashboard receives progress over a WebSocket. But a sequence of events is not a complete record of the job.
A merchant might close the tab during indexing and return later. On connection or reconnection, the dashboard needs the current state, not just future notifications.
I also keep polling as a fallback for environments where the socket cannot connect. Losing a progress connection should not automatically mean the ingestion failed.
Make completion safe to retry
A background job can finish successfully while the request acknowledging that success fails. Retrying the completion call must not charge the customer again.
The settlement path therefore needs to be idempotent. Duplicate completion calls should preserve the same completed job and charge, rather than create another billing event.
Development and production also need separate storage and queues. Otherwise an upload can trigger a consumer connected to a database where that job does not exist.
What I would carry into another RAG project
Use stable document identities. Make the handover from browser to background processing explicit. Keep job state outside the browser. Treat notifications as a convenience over authoritative state. Design retries and settlement together.
Those choices make the pipeline easier to reason about when uploads, indexing, network connections, and billing do not all succeed at the same time.
I’m building this in ProxyAI, an AI customer-chat platform that is free to start, with no subscription fees and pay-as-you-go AI usage.
If you have built a similar pipeline, how do you handle replacing documents without leaving stale content available during reindexing?
Top comments (0)