DEV Community

Checkout Flow Digital
Checkout Flow Digital

Posted on Originally published at checkoutflowdigital.com

Webhook vs API: Choosing the Right Integration Pattern

"Should this integration use the API or a webhook?" comes up at the start of almost every integration project, and it is slightly the wrong question. Webhooks and APIs are not alternatives. They cover two different halves of the same problem:

  • An API is how your system requests data from another system, or instructs it to do something.
  • A webhook is how the other system lets yours know that something just happened.

Integrations that survive production usually use both. This post covers the difference, when each one fits, and the handful of patterns (signature verification, idempotency, retries, reconciliation) that make the combination reliable. The examples come from e-commerce, Shopify in particular, but the same patterns apply to payment providers, CRMs and most SaaS platforms.

Push vs pull

PULL — the API                          PUSH — the webhook

your service                            platform
    │  GET orders updated since 10:00       │  POST /webhooks/orders  (event)
    ▼                                       ▼
platform                                your endpoint
    │  200 + data                           │  200 OK  (acknowledgement)
    ▼                                       ▼
your service                            platform

You decide when to ask.                 The platform decides when to tell you.
Enter fullscreen mode Exit fullscreen mode
  • With pull, you control timing, volume and exactly what you fetch. The cost is latency (you only learn about a change when you next ask) and wasted calls (most polls return nothing new).
  • With push, you hear about an event seconds after it happens, without polling. The cost is a public endpoint that has to be fast, secure, and tolerant of duplicates and gaps.

Webhook vs API at a glance

API (pull) Webhook (push)
Who initiates Your system The other platform
When data arrives Whenever you ask: on demand or on a timer Usually within seconds of the event
What it can do Read and write Notify only; you still call the API to act
Good at Creating and updating records, searches, backfills, catching up after an outage Reacting to events: payment captured, order created, stock changed
Typical failure Polling so often you hit rate limits, or so rarely the data goes stale Deliveries that are missed, late, duplicated, out of order or forged
What you have to build Authentication, pagination, rate-limit handling, version upgrades A public HTTPS endpoint, signature verification, de-duplication, a queue

A webhook is sometimes called a "reverse API": same building blocks (HTTP, JSON), opposite direction. It rarely replaces the API, though. A webhook payload tells you that something happened; reading the full record or changing anything usually still goes through the API.

A real flow: paid order → accounting software

Say every paid Shopify order must show up as a sale in an accounting tool. A design that holds up in production looks like this:

REAL-TIME PATH
Shopify ──(webhook: order paid)──▶ your endpoint
  1. verify the signature on the raw body     → 401 if invalid
  2. record the delivery id; seen it before?  → 200, do nothing
  3. put a job on a queue                     → 200 immediately
                     │
                     ▼
worker
  4. read the full order through the Admin API (if the payload isn't enough)
  5. sale already created for this order? skip · otherwise create it
     through the accounting API
  6. transient error: retry with backoff · permanent error: stop and alert

SAFETY NET (hourly or daily)
scheduler ──▶ Admin API: "orders updated since <last run minus an overlap>"
  7. hand each order to the same worker: step 5's order-level check skips
     orders a webhook already synced (step 2 can't: an order fetched from
     the API carries no webhook delivery id)
Enter fullscreen mode Exit fullscreen mode

Speed comes from the webhook; control and recovery come from the API. Integrations tend to break exactly where a team relied on only one of them: polling-only setups are slow and eat the rate limit, and webhook-only setups quietly lose data the first time the endpoint is down.

Webhooks: what you have to get right

1. Verify every delivery before trusting it

Your endpoint is a public URL, so anyone can POST to it. Shopify signs each HTTPS delivery: the X-Shopify-Hmac-Sha256 header holds a base64-encoded HMAC-SHA256 of the raw request body, keyed with your app's client secret. Recompute it and reject anything that doesn't match.

const crypto = require('node:crypto');

function isValidShopifyWebhook(rawBody, hmacHeader, secret) {
  if (typeof hmacHeader !== 'string' || hmacHeader.length === 0) return false;
  const expected = Buffer.from(
    crypto.createHmac('sha256', secret).update(rawBody).digest('base64'),
  );
  const received = Buffer.from(hmacHeader);
  // timingSafeEqual throws when lengths differ, so check that first
  return expected.length === received.length && crypto.timingSafeEqual(expected, received);
}
Enter fullscreen mode Exit fullscreen mode

Two mistakes cause most "my HMAC never matches" bugs:

  • Hashing a parsed and re-serialized body. If your framework parses JSON before your handler runs, whitespace and key order change and the signature will never match. Capture the raw bytes.
  • Comparing with ===. Use a constant-time comparison (crypto.timingSafeEqual in Node, hmac.compare_digest in Python).

Most payment providers sign webhooks the same way (an HMAC over the raw body, sent in a header). What changes is the header name, sometimes hex instead of base64, and sometimes a timestamp included in the signed content to limit replays.

2. Acknowledge fast, process later

Senders give you a short window. Shopify, for example, documents a one-second connection timeout and five seconds for the whole request. It treats any response outside the 2xx range (redirects included) as a failure and retries failed deliveries up to 8 times over 4 hours with exponential backoff. If failures persist, the subscription can be removed.

So the request handler should do the minimum (verify, record, enqueue) and return 2xx. Calls to your ERP, accounting tool or email provider belong in a worker, where a slow third party can't make the sender give up on you.

3. Expect duplicates and disorder: make processing idempotent

Webhooks are delivered at least once. The same event can arrive twice (a timeout on your side, then a retry on theirs), and events can show up in a different sequence than they occurred.

The usual fix is an idempotency ledger: claim a key before doing any work, and skip the delivery if the key is already taken.

on delivery:
  key = delivery id
        (fallback: topic + resource id + updated_at)
  if not ledger.claim(key):   # atomic insert against a unique constraint
      return 200              # duplicate: acknowledge, do nothing
  enqueue(job)                # if this fails: release the claim, return 5xx
  return 200
Enter fullscreen mode Exit fullscreen mode

For Shopify, the documented header for de-duplicating deliveries is X-Shopify-Webhook-Id. X-Shopify-Event-Id is shared by deliveries to different subscriptions that come from the same merchant action, which makes it useful for correlation.

A few details matter more than they look:

  • The ledger must be shared by every instance that handles webhooks: a table with a unique constraint, or Redis SET key value NX EX <ttl>. An in-memory set only works on a single process.
  • Keep delivery keys long enough to cover the sender's retry window.
  • Ordering: don't assume it. When order matters, compare the resource's updated_at with what you last stored and ignore stale updates, or re-read the current state from the API.
  • Business-level idempotency: the delivery ledger only absorbs repeated deliveries. The step that creates the invoice needs its own guard at the order level: check whether a sale already exists for this order (in your own records, or by external reference in the target system), or pass an idempotency key if the target API supports one. Make that check atomic, for example with a unique constraint on the order id in your sync records, so two workers can't both pass it at the same time. That order-level check is also what makes reconciliation safe (section 5).

4. Retry with backoff, and know when to stop

Inside the worker, the downstream API will fail from time to time. Treat the two kinds of failure differently:

  • Transient (network error, timeout, 429, 5xx): retry with exponential backoff and jitter, and honour Retry-After when the API sends it.
  • Permanent (400 validation error, 404, a business rule): stop, record the failure, alert a human. Retrying will not help.
// "Full jitter": a random delay between 0 and the exponential cap
const delayMs = Math.random() * Math.min(60_000, 500 * 2 ** (attempt - 1));
Enter fullscreen mode Exit fullscreen mode

Jitter matters more than it seems. After an outage, thousands of jobs retrying on the same schedule hit the recovering API at the same moment. Jobs that fail for good should land somewhere a person actually looks.

5. Never rely on webhooks alone: reconcile on a schedule

Endpoints go down during deploys, certificates expire, a bug returns 500 for an hour, a retry window runs out. Sooner or later a webhook will be missed.

Reconciliation is a scheduled job, hourly or daily, that asks the API what changed since the last run (with some overlap) and hands every result to the same worker the webhooks feed. With Shopify, that means a GraphQL Admin API query on orders filtered by updated_at, sorted by UPDATED_AT and paginated with cursors.

Don't count on the webhook delivery ledger to prevent duplicates here. An order fetched from the API carries no webhook delivery id, so a ledger keyed on delivery ids cannot tell that a webhook already handled it. What prevents a second sale is the order-level check inside the worker: has a sale already been created for this order? Orders a webhook already synced stop there, and the missed ones get processed.

This is the piece that turns "usually works" into "trustworthy".

APIs: what you have to get right

Stay inside the rate limits

Every API meters its usage. Shopify's GraphQL Admin API, for instance, charges a calculated cost per query against a bucket that refills continuously, and the refill rate depends on the store's plan. Request only the fields you need, paginate, and back off when you're throttled (an HTTP 429, or a THROTTLED error in a GraphQL response) instead of hammering.

Keep credentials on the server

Access tokens and API secrets belong on a server or in a secret manager. Never put them in a storefront theme, in front-end JavaScript or in a repository. Ask for the minimum scopes: a reconciliation job that reads orders needs read_orders, not write access to everything.

Pin versions and plan upgrades

APIs are versioned, and so are webhook payloads (Shopify sends the version in an X-Shopify-API-Version header). Pin a version explicitly, keep an eye on deprecation announcements, and run your tests against a new version before switching.

Failure modes and what handles them

What goes wrong What protects you
Forged request to your endpoint HMAC verification on the raw body, then 401
Same event delivered twice Idempotency ledger keyed on the delivery id
Events arrive out of order Compare updated_at, or re-read current state through the API
Slow handler, so the sender times out and retries Acknowledge fast, process in a queue
Downstream API flaky or rate-limited Retries with exponential backoff, jitter and Retry-After
Downstream rejects the data Permanent-error path: stop, log, alert
Endpoint down for longer than the retry window Scheduled reconciliation through the API
Worker crashes after you already returned 2xx A durable queue (database table, SQS, Cloud Tasks…), not process memory

Common mistakes

  1. Doing the real work inside the webhook request, then timing out.
  2. Verifying the signature against re-serialized JSON instead of the raw body.
  3. De-duplicating with an in-memory set on a multi-instance deployment.
  4. Returning 2xx while the job only lives in process memory.
  5. Polling every few seconds "to be safe" and burning through the rate limit.
  6. Webhook-only integrations with no reconciliation job.
  7. Shipping an Admin API token in theme or front-end code.
  8. Retrying a 400 forever.

Decision checklist

Which one?

  • Something just happened and you must react in seconds (payment confirmed, order created): a webhook to hear about it, the API to act on it.
  • You're writing to another system (creating or updating records): API.
  • Batch work such as a nightly export, a report or a backfill: scheduled API calls.
  • No webhook exists for the event you care about: poll the API on a sensible interval, inside the rate limits.
  • Money or inventory is at stake: both, with reconciliation on top.

Before a webhook endpoint goes live:

  • Signature verified on the raw body with a constant-time comparison
  • 2xx returned within a couple of seconds; slow work handed to a durable queue
  • Idempotency ledger shared by every instance
  • Retries with backoff and jitter; permanent errors reach a human
  • Scheduled reconciliation that overlaps the previous run and feeds the same worker
  • An atomic, order-level "sale already created?" check in that worker (for example, a unique constraint on the order id), so webhooks and reconciliation can't both create the sale
  • Secrets in the environment or a secret manager; minimal API scopes
  • One log line per delivery: accepted, duplicate, rejected or failed

FAQ

Is polling an API bad practice?

No. Polling on a sensible schedule that stays under the rate limits is how you catch up after gaps, and it's the only option when a platform has no webhooks. Where it falls short is reacting in near real time: that's the webhook's job.

Can I skip the queue at low volume?

If your handler only writes to your own database and finishes well inside the sender's timeout, you can start without one. Once a third-party API call sits in the request path, though, your acknowledgement depends on someone else's latency, and that is when duplicates and dropped events start.

Companion code

If you'd like to see these patterns as runnable code, there is a companion repository: checkoutflowdigital/shopify-webhook-patterns. It holds small, dependency-free examples in Node.js (18+) and Python (3.10+):

  • HMAC signature verification (raw body, constant-time comparison)
  • an idempotency ledger keyed on X-Shopify-Webhook-Id, with a run-once helper
  • retries with full-jitter exponential backoff and a permanent-error path
  • a scheduled reconciliation job against the Shopify GraphQL Admin API (pagination, throttling), in Node.js only

The first three exist in both languages and come with tests. The reconciliation job is a Node.js sketch that leaves the order-level "already synced?" check to the handler you pass in.

It is reference code to read and adapt, not an official Shopify library and not a drop-in package. Swap the in-memory store and queue for your own database and queue before taking the ideas to production.


This article is adapted from Webhook vs API: what's the difference, with e-commerce examples, first published by Checkout Flow Digital — Commerce & Payment Integration.

Top comments (1)

Collapse
 
challan116ux profile image
challan116-ux •

The raw-body HMAC point deserves emphasis — hashing re-serialized JSON is still the #1 'signature never matches' bug we see, ahead of even wrong secrets. Agree on ack-fast too: verify, persist, enqueue, 2xx, and let the worker own retries, otherwise a slow downstream turns into duplicate deliveries. The scheduled reconciliation overlap is what makes the whole thing trustworthy once the sender's retry window expires.