Disclosure: I run NxFlowAI, a custom AI workflow agency. The patterns below are tool-neutral.
Before you wire an LLM into production messaging, make the handoffs on both sides of it boring: inbound events processed once, outbound replies sent once, and partial writes impossible to miss. That is the unglamorous work a custom AI workflow agency that designs retry-safe handoffs should finish before anyone tunes a prompt.
Three failures that look like "the AI is wrong"
- A customer gets the same reply twice because the send step timed out and the job runner retried it.
- The CRM shows a lead with no owner and no phone number, while the bot already told the customer "all set".
- The model is called twice for the same message, so two different drafts reach the approval queue.
None of these are model problems. They are handoff problems, and adding an LLM makes them more expensive because every retry now costs tokens and can produce a different answer.
Inbound: record the event before doing anything
def handle_inbound(event):
key = f"{event.channel}:{event.provider_message_id}"
if not processed.insert_if_absent(key, received_at=now()):
return "duplicate" # already handled, do nothing
enqueue("extract_and_route", key)
The insert-if-absent must be atomic (a unique index is enough). Do the slow work, including the LLM call, in the queued job, not in the webhook handler.
Outbound: an outbox, not a direct send
def queue_reply(conversation_id, draft_id, text):
outbox.insert(id=f"{conversation_id}:{draft_id}", text=text, status="pending")
def deliver(row):
if row.status != "pending":
return
result = provider.send(row.text, client_ref=row.id) # pass your own reference if supported
outbox.update(row.id, status="sent", provider_id=result.id)
The outbox row id is derived from the approved draft, so approving twice or retrying delivery cannot create a second message. If your provider accepts a client reference, store it; it makes reconciliation much easier.
Partial writes: make "half done" visible
Write a single handoff_status field on the CRM record: received, matched, owned, replied. A lead stuck at matched for an hour is an alert, not a surprise next week.
Where the LLM goes
Only after these three exist: inside the queued job, with its output stored against the event key. If the job retries, it reuses the stored draft instead of generating a new one.
Checklist before go-live
- Unique key per inbound event.
- Outbox with deterministic ids.
- One status field that shows where a handoff stopped.
- A dead-letter list for events that fail repeatedly, reviewed by a person.
- Stored LLM outputs, reused on retry.
In our 72-hour audits, this list usually generates more findings than the prompts do.
Top comments (0)