DEV Community

Nabeel Hassan
Nabeel Hassan

Posted on

A Webhook Tells You Something Happened, Not That You Have Everything

I build VoiceDash, a white-label portal where agencies show their own clients the calls their AI voice agents handled: transcripts, recordings, analytics, all under the agency's brand. The product has exactly one job that the client actually notices. When they open the dashboard, every call should be there.

For a long time, my answer to "how do calls get in?" was "the webhook". Retell posts a call event, I store a row, the dashboard reads the table. There was also a Sync Calls button for the rare case where something got missed.

Last week I spent a day admitting that this design was backwards. A webhook tells you something happened. It does not tell you that you have everything. Those are different promises, and only the second one is what my users care about.

How "the webhook" quietly failed

Nothing caught fire. The failures were all quiet:

  1. A new customer onboarded to an empty dashboard. Connecting a Retell agent imported nothing. The webhook only fires for calls made after you register it, so an agency with months of history saw zero calls until someone opened the agent's page, and even then the import stopped at the latest 100.
  2. Gaps never healed. The sync button read the newest page and stopped at the first page with nothing new. If calls went missing further back (say, while a webhook URL was broken), nothing would ever go back for them.
  3. A page left open went stale. The dashboard loaded once. New calls only appeared on refresh.
  4. One Retell agent, two workspaces. A call was stored once, globally unique by Retell's call id, for whichever VoiceDash agent saw it first. Connect the same Retell agent in a second workspace (a teammate, a re-onboarding) and that workspace only ever saw calls from after it took over the webhook.
  5. Calls from today that were not from today. Calls that never connected have no start timestamp. My importer stamped them with the import time, so they showed up as calls made a minute ago.

Every one of these came from the same assumption: that the webhook stream is the source of truth and everything else is a patch. Flip it, and most of them disappear.

The source of truth is the provider's list endpoint

Retell has a paginated list-calls API. That is the ledger. The webhook is a fast-path notification that lets a call appear within seconds, but correctness has to come from reconciling against the list.

The naive version of reconciliation is "re-read everything every time", which is fine for an agent with 40 calls and terrible for one with 40,000. So each agent now carries one extra column:

callsSyncedUntil      DateTime?
Enter fullscreen mode Exit fullscreen mode

It means: every call that started before this moment is already stored. A sync reads calls oldest first from just behind that point and moves it forward after every page.

// Everything up to the newest call on this page has now been read, in
// order, from a point that was already complete.
const starts = calls.map((c) => c?.start_timestamp).filter((t): t is number => typeof t === "number");
if (starts.length > 0) await advanceWatermark(opts.agentId, new Date(Math.max(...starts)));
Enter fullscreen mode Exit fullscreen mode

The ordering is the whole trick. If you read newest first, a sync that dies on page three leaves you with a hole you cannot describe: you have the newest 300 calls and no idea where the gap starts. Read oldest first from a point that was already complete, and after every page the claim "everything before X is stored" is still true. An interrupted sync is not a failure, it is just a sync that has not finished, and the next one resumes where it stopped.

The one exception is an agent with no watermark at all, a fresh connection. There I read a single page newest first before switching to ascending, purely so the dashboard fills in immediately instead of starting with the oldest calls it has.

Two details that make the watermark honest

It only moves forward. Two syncs for the same agent can overlap: a page open in two tabs, the onboarding import still running in the background. So the update is conditional:

async function advanceWatermark(agentId: string, until: Date) {
  // Only ever moves forward, including when two syncs of one agent overlap.
  await prisma.agent.updateMany({
    where: { id: agentId, OR: [{ callsSyncedUntil: null }, { callsSyncedUntil: { lt: until } }] },
    data: { callsSyncedUntil: until },
  });
}
Enter fullscreen mode Exit fullscreen mode

No lock needed: a slower sync finishing late cannot drag the watermark backwards.

Every sync re-reads a little behind it. A call that is still in progress, or that Retell lists a bit late, could otherwise fall just before the watermark and never be read. So the read starts two hours early:

const OVERLAP_MS = 2 * 60 * 60 * 1000;
Enter fullscreen mode Exit fullscreen mode

Re-reading is cheap because storing is idempotent. Calls the agent already has are skipped, with one exception: a row stored mid-call (duration still zero) gets its final duration, transcript and recording when the finished version shows up. A routine sync is still one request.

Time budgets instead of hoping

Importing a long history inside an HTTP request is how you get a timeout halfway through onboarding. So the sync takes a budget and returns where it stopped:

export interface SyncResult {
  /** Calls newly added to this agent. */
  imported: number;
  /** False when the time budget ran out before every call was read. */
  complete: boolean;
  /** Where to continue when not complete. */
  resume?: SyncResume;
}
Enter fullscreen mode Exit fullscreen mode

Connecting an agent now imports for up to 20 seconds before responding, and the onboarding confirmation says how many past calls came in. If the history is longer than that, the rest finishes in the background with Next.js after(), using the resume cursor. Because the watermark advances per page, even a background run that gets killed leaves the next sync a correct place to start.

I also changed Sync Calls to report calls it actually added ("Imported 3 new calls", or "No new calls") instead of everything Retell returned, most of which was already there.

Nobody should have to press Sync

Once a routine sync costs one request, there is no reason to make a human trigger it. The analytics and conversations pages now use a small hook:

sync({ evenIfHidden: true });
const timer = setInterval(() => sync(), SYNC_EVERY_MS);
document.addEventListener("visibilitychange", onVisibilityChange);
Enter fullscreen mode Exit fullscreen mode

It syncs when the page opens, every minute while the tab is visible, and when the tab comes back into view (unless the last sync was under 15 seconds ago). A running flag stops overlapping requests from the same tab.

The evenIfHidden on the first call came from a follow-up commit the same evening. I had originally gated every sync on document.visibilityState === "visible", which meant a dashboard opened in a background tab, a middle click or a restored session, showed stored calls and skipped its opening sync until someone switched to it. Opening a page should always sync once. Only the repeat should wait for visibility.

After each sync the page refreshes in place: numbers update without loading placeholders, new calls join the list without losing loaded pages or the conversation you have open. The button is still there as a fallback that re-reads the whole history, but it is no longer part of the normal path.

The uniqueness constraint was modelling the wrong thing

The multi-workspace bug was not a sync bug. It was a schema bug. platformCallId was @unique, which encodes "a call exists once in VoiceDash". The real rule is "a call exists once per agent", because each workspace is a separate customer with its own view.

// One copy of a call per agent: a Retell agent connected in several
// workspaces shows its calls in every one of them.
@@unique([agentId, platformCallId])
@@index([platformCallId])
Enter fullscreen mode Exit fullscreen mode

Retell still sends webhooks to only one address, the most recently connected workspace. The others get each call from their own automatic sync, through their own API key, which can only list calls its own Retell account owns. That last part matters: copying rows across workspaces would have been easier, and it would have been a tenant isolation bug.

Two smaller pieces came with it. Deleting the agent that owns the webhook now hands it to the newest remaining agent for the same Retell agent, but only if the webhook still points at the deleted one, so a webhook someone has set up elsewhere since is left alone. And the migration ran in two steps around the deploy: new indexes first, then dropping the old unique index once no running code depended on it. It also reset every watermark to null, because calls skipped earlier ("another workspace already has this") now belonged to every agent and needed reading again. The watermark design made that a one-line UPDATE instead of a backfill script.

What I would tell myself before the first webhook

  • Decide which source is the ledger. For third-party events it is almost always the provider's list API. Webhooks make things fast; reconciliation makes them correct.
  • Store progress as a claim you can keep true. "Everything before X is stored" survives crashes if you read in order and advance after each page. "I synced at 3pm" does not.
  • Let overlap be cheap. If writes are idempotent, re-reading a window behind your watermark costs almost nothing and covers late and unfinished records.
  • Check your unique constraints against tenancy. A globally unique external id in a multi-tenant table is a decision about who owns the data, whether you meant it or not.

I wrote earlier about having five code paths write the same subscription row on purpose. This is the same instinct from the other side: do not trust any single delivery path to be complete, and make the convergent path cheap enough to run all the time.

If you are building client portals on top of Retell, this is what VoiceDash does under the hood: voice-dash.com.

Top comments (3)

Collapse
 
beusebiu profile image
Eusebiu Balan •

Billing pushed me to the same place. Lemon Squeezy webhooks are the fast path, and every few minutes a job asks their API again about accounts whose local state looks wrong: a subscription id sitting on a free account, or a subscription that expired in the last day.

Much cheaper than walking the full list, though it only catches the gaps you can see in your own data.

Collapse
 
jeemmo profile image
Azeem Javed •

This matches what I landed on for company-filing alerts: a live stream for speed, plus a daily job that re-checks every watched company and catches whatever the stream missed during reconnects or outages. One detail worth stealing: the first time the daily job sees an item, record it as a silent baseline instead of alerting, so a backfill or a newly added account doesn't flood users with old events.

Collapse
 
elijahbrown profile image
Elijah Brown •

Filling the first sync during onboarding so the dashboard is not empty is the useful product fix. When you invent demo agency and client fixtures for that path, use reserved fiction phones (US 555-0100 to 555-0199, UK 020 7946 0xxx) and domains you own, so a mis-scoped export during a workspace handoff cannot publish a real contact.