DEV Community

Corsair
Corsair

Posted on

How to Prevent Infinite Loops in Bidirectional API Syncs


A sales rep updates a phone number in the CRM. Thirty seconds later, the same field has been rewritten four times, three systems are stuck firing events at each other, and nobody on the team can say why the sync queue keeps climbing. If you have built or maintained a bidirectional integration between two systems of record, this probably looks familiar.

Bidirectional API syncs are one of the more deceptively hard problems in API integration. Two systems, each with their own webhooks, each capable of writing back to the other, sound simple on a whiteboard. In production, a single genuine edit can spiral into dozens of redundant writes, wasted API calls, and in the worst cases a loop that never settles on its own.

This post walks through why infinite loops happen in bidirectional syncs, why deduplication and retry logic alone will not stop them, and the change ownership, origin metadata, and value comparison strategies that do. These are the same best practices for building reliable API integrations more broadly, not just a fix for this one failure mode. Whether you are designing a new API and data integration from scratch or debugging one that already misbehaves, the goal by the end of this is the same: writes that happen once, settle, and stop.

How One CRM Update Can Spiral Into an Infinite Sync Loop

Picture two systems kept in sync by a middle layer, call it System A (a CRM) and System B (an ERP), connected through a sync engine that listens for change events on both sides. Here is the exact chain that turns one legitimate edit into a loop.

  1. A user changes a phone number in System A. That edit fires a change data capture event, often shortened to CDC, which the sync engine picks up as a new fact to propagate.
  2. The sync engine calls System B's API and writes the new phone number.
  3. System B commits the write. Because System B has its own event pipeline (a webhook, a CDC stream, an audit log listener), that commit emits a change event of its own. System B has no way of knowing the change originated from an integration rather than a human.
  4. The sync engine, listening to System B's events the same way it listens to System A's, sees this new event and treats it as a fresh instruction: propagate the change back to System A.
  5. System A receives the write, commits it, and its own event pipeline fires again.
  6. The loop repeats, sometimes settling after a couple of bounces, sometimes continuing indefinitely if formatting or enrichment logic changes the value slightly on every pass.

The root cause is rarely a broken API call or a flaky network. It is that nothing in the pipeline checks whether an incoming change actually originated from the integration itself. Every event, whether caused by a human typing into a form or by the sync engine's own previous write, looks identical by the time it reaches the listener. Without a way to tell those two cases apart, the sync engine cannot help but treat its own echo as new work.

Duplicate Delivery, Retried Writes, and Echo Events Aren't the Same Failure

It is tempting to lump every repeated event under one label and reach for one fix. Three distinct failure modes get confused with each other constantly, and only one of them causes infinite loops.

  • Duplicate delivery: the same event, carrying the same transaction ID, redelivered by a message broker or webhook provider that guarantees at least once delivery rather than exactly once. Nothing in the underlying system has changed. Track the event ID you have already processed and drop anything you see twice.
  • Retried write: your own sync engine resending an HTTP call because the first attempt timed out or returned a retryable error. The destination may have actually applied the first write, so a naive retry risks a second one. Conditional headers such as ETag checks, along with idempotency keys on the write endpoint itself, catch this cleanly by making the second attempt harmless when the state has not moved.
  • Echo event: a genuinely new, freshly committed transaction on the target system, triggered by the sync engine's own write. It has a new transaction ID, a new timestamp, and a new payload. Deduplication will not catch it, because it is not a duplicate. Idempotency keys will not catch it, because it is not a retried write.

This distinction matters because teams often spend weeks hardening how incoming webhook events are routed and verified or adding retry logic with backoff, watch the duplicate and retry problems disappear, and are then confused when the sync still loops. Those fixes were solving a different problem. Echo events need a different kind of defense, covered next.

Assigning Change Ownership When Systems Share a Record

One of the most effective structural fixes has nothing to do with detecting echoes after the fact. It is deciding, up front, which system is allowed to change which fields, so an incoming event that touches a field it does not own can be ignored outright.

Whole record ownership, where one system is declared the single source of truth for an entire object, sounds clean but rarely survives real data. A CRM and an ERP both have legitimate reasons to touch a shared customer record: the CRM owns the sales relationship (name, phone, lifecycle stage), while the ERP owns anything downstream of a signed contract (billing address, payment terms, tax ID). Forcing one system to own the whole record either strips the other of fields it genuinely needs to edit, or invites the exact write back that starts the loop in section one.

Field level ownership is more precise. Instead of asking which system owns this record, you ask which system owns this field, and only propagate a change when the field being changed actually belongs to the receiving system's list. In practice this is implemented with a changedFields array attached to every event, listing exactly which fields were touched rather than sending the full record on every change.

{
  "recordId": "cust_48213",
  "changedFields": ["phone", "lifecycleStage"],
  "source": "crm"
}
Enter fullscreen mode Exit fullscreen mode

Before the sync engine writes anything to System B, it checks changedFields against the fields System B is allowed to own. If none match, there is nothing to propagate, and the event is dropped without a single API call. A write that never happens cannot trigger an echo.

4. What Origin Tracking Metadata Can (and Can't) Tell You

Field ownership handles the case where systems cleanly divide responsibility. It does not help when both systems legitimately need to write to the same field, which is where origin tracking comes in.

The idea behind a changeOrigin field is simple: every write your sync engine makes carries a marker identifying who made it, typically a client ID, an integration version, and sometimes a request ID.

{
  "recordId": "cust_48213",
  "field": "phone",
  "value": "+1 555 0134",
  "changeOrigin": {
    "client": "sync_engine_v3",
    "actor": "integration"
  }
}
Enter fullscreen mode Exit fullscreen mode

When the sync engine later receives an event and sees a changeOrigin matching its own client ID, it can safely conclude this is its own echo and stop, rather than propagating the change back to where it came from. In theory this closes the loop cleanly. In practice it has three limitations.

  • Often blank: many APIs only populate an origin or actor field when the calling client explicitly sets a custom header on the write, and it is easy to miss that step on endpoints the sync engine calls infrequently.
  • Gets stripped: some destination systems do not store arbitrary metadata on the record, so a write that includes changeOrigin on the way in comes back out through the webhook or CDC stream with no trace of it.
  • Can't be the sole defense: it is a fast, cheap first check, not something you can rely on for every provider, every event type, and every edge case.

That last point raises an obvious question: what do you do when the destination does not preserve origin metadata at all.

Catching Your Own Echo When the Destination Doesn't Preserve Origin Metadata

When origin metadata is missing or unreliable, the fallback is to compare values instead of trusting labels. Before writing an incoming change to System A, the sync engine checks whether the incoming value actually differs from what System A currently holds. If the value is already +1 555 0134 and the incoming echo says +1 555 0134, there is nothing to write, so the sync engine moves on without an API call, and never generates a new event to feed back into System B.

This sounds straightforward but the implementation needs care in two places.

  • Normalize before comparing: "+1 555 0134", "1 555 0134", and "(555) 013 4" are the same value to a human and different strings to a naive equality check. Dates, currency amounts, and whitespace need the same treatment, or you end up writing back values that are functionally identical but fail the comparison.
  • Watch the timing: a value check has to avoid silently swallowing a legitimate concurrent edit. If a second user changes the same field microseconds after the echo arrives, comparing against a value cached earlier in the pipeline can drop a change that should have gone through. Read the current value as close to the write as your database or API allows.

It is worth restating why deduplication and idempotency keys, the tools from section two, do not solve this on their own. Every echo is a fresh, validly committed transaction with its own transaction ID. It is not a replay and not a retried write, so none of that machinery ever fires. Value comparison sits closer to the write itself, and it is the layer that actually intercepts an echo once metadata has already failed to.

If you are building this on an integration layer that exposes a hook before a webhook driven write goes through, this comparison is the natural place to put it. A before hook that can inspect the payload and skip processing entirely means the no effect decision happens in one place, next to the write itself, instead of being scattered across every consumer of the event. Storing the current value somewhere queryable, such as a synced entity table kept fresh by every API call and webhook, is what makes reading the current value right before deciding practical instead of theoretical.

Handling Concurrent Edits and Verifying Systems Actually Reach a Stable State

Everything so far assumes one change moving through the pipeline at a time. Real systems have two people editing the same record within seconds of each other, and the sync engine needs a rule for who wins that does not also reintroduce a loop. Three approaches handle this, and each one catches a different failure while leaving a different gap.

  • Ownership boundaries: the field level approach from section three prevents the conflict from existing at all for any field only one system can touch. This is the strongest guarantee available, but it says nothing about a field like status or notes that both systems have a legitimate reason to write.
  • Version or ETag checks: every write includes the version number it last read. If System B tries to write using version 3 but the current version is already 4, the write is rejected and the sync engine reads the latest state again before retrying. This catches real conflicts precisely, but it depends entirely on the destination supporting and returning version numbers, which not every API does.
  • Timestamp based decisions (last write wins): compare when each change happened and let the most recent one stand. This is the easiest to implement and works with almost any system, since most APIs return an updated timestamp by default. Its weakness is clock skew: if the two systems' clocks are not tightly synchronized, most recent can be wrong, and a genuinely newer change can lose to an older one that simply arrived with a later local timestamp.

Most reliable syncs combine these rather than picking one: ownership boundaries remove as many fields from the conflict question as possible, version checks handle the fields both systems can touch wherever supported, and timestamp comparison is the fallback everywhere else.

None of this matters if you cannot verify it works, which is where stable state testing comes in. Make one legitimate change in System A and watch what happens across both systems afterward.

  • A stable integration propagates the change to System B exactly once, and after that single propagation, no further events fire on either side.
  • An unstable integration keeps generating events: the value bounces between two states, or the same value gets written repeatedly with no functional change, or event counts on your monitoring dashboard never return to zero after one edit.

Running this test deliberately, for every field and every direction, before a bidirectional sync goes into production, catches loops in staging instead of in a support ticket.

Every pattern covered here (change ownership, origin metadata, value comparison, and conflict resolution) comes down to the same discipline: know whether a change is genuinely new before you act on it. That discipline is exactly what an integration layer should handle for you, rather than something every team rebuilds inside its own sync engine. Corsair is built as that layer, with webhook routing, retry handling, and a database that stays fresh on every API call and webhook as part of the SDK itself, so the logic above sits on top of a foundation instead of replacing one you have not built yet. If your team is wiring up a new API integration or hardening a bidirectional sync that already misbehaves, it is worth seeing how much of this Corsair already handles.

FAQs

What is the difference between an infinite sync loop and normal duplicate webhook delivery?

Duplicate delivery means the same event, with the same transaction ID, arrives more than once, with no new state change, and is solved by tracking event IDs and dropping repeats. An infinite loop is caused by echo events, which are new, validly committed transactions triggered by the sync engine's own earlier write. Because each one has a new transaction ID, event ID based deduplication never recognizes them as anything other than fresh work.

Can idempotency keys alone prevent bidirectional sync loops?

No. Idempotency keys protect against retried writes, where your own sync engine resends the same request after a timeout, making sure a retry does not create a second write. They do nothing for echo events, since an echo is not a retry of a request you sent, it is a brand new event coming from the destination in response to a write that already succeeded. Preventing loops needs origin tracking or value comparison on top of idempotency, not instead of it.

How do you decide which system owns a shared field like phone number or billing address?

Assign ownership at the field level rather than the whole record. Pick the system where the field is edited in the normal course of work, for example the CRM for contact details and the ERP for billing information, and only allow writes to that field from that system. Attach a changedFields array to every event so the receiving system can check whether an incoming change actually touches a field it owns, and skip it entirely if it does not.

What should you do when a destination API does not return version numbers or ETags?

Fall back to value comparison. Read the current value of the field immediately before deciding whether to write, normalize both values so equivalent values compare as equal, and skip the write if nothing has actually changed. This will not resolve every concurrent edit as precisely as version checks would, but it stops the specific failure that causes loops, which is writing back a value that never actually changed.

How do you test whether a bidirectional integration is actually stable before shipping it?

Make one legitimate change on one side and watch both systems' event activity afterward. A stable integration propagates that change exactly once and then goes quiet. If you see the value bounce, or the same write repeat with no functional change, or event counts that never return to zero, the integration is not stable yet. Running this check for every field and direction in staging, before production traffic hits it, is far cheaper than debugging a live loop.

Top comments (0)