The server reboots at two on a Saturday morning for a security update, which is what it's supposed to do. The app doesn't come back, because nobody told the process manager to start it on boot.
It's a fundraising weekend. Orders keep arriving, Shopify keeps sending webhooks into nothing, and after four hours of retries it stops.
On Monday the app is running again and looks perfectly healthy. It's missing two days of sales, and every seller's total is short.
The two rules, and how they interact
An app has five seconds to respond to a webhook. Miss it and Shopify retries — up to eight times over four hours — and if failures persist, the subscription is removed. Shopify's guidance for recovering is to re-subscribe and import the missing data.
Separately, a webhook can arrive more than once after a timeout or a retry.
Those two rules interact in a way that catches self-hosted apps specifically. A small server under load, writing to its ledger before it answers, can take longer than five seconds. Shopify counts that as a failure and sends the webhook again — so the app records the same order twice. Not because anything was wrong with the delivery. Because it was slow to say thank you.
The fix is ordering rather than speed: acknowledge first, do the work afterwards, and make the work safe to repeat.
Read together, those rules say something that should change how you design the thing: webhooks are a notification that something happened, not a guarantee you'll hear about everything. An app whose records are built only from webhooks is complete exactly as long as it has never been down.
Every self-hosted app is eventually down.
Reconcile against the orders
So the app doesn't trust its own memory. On a schedule — daily is enough for most — it fetches the orders Shopify has for a window, compares them against what it recorded, and processes the gaps.
ts
`/** An order as fetched from the Admin API, with the seller attribute it was placed with. */
type ShopifyOrder = { id: string; createdAt: string; sellerRef: string | null };
/** What the app recorded when a webhook arrived. */
type Recorded = { orderId: string; webhookId: string };
type Finding =
| { kind: "missed"; orderId: string; sellerRef: string }
| { kind: "unattributed"; orderId: string };
/**
- Compare what Shopify says happened in a window with what the app heard about. *
- Webhooks are the fast path, not the record. The attribute rides on the order
- itself, so anything the app missed while it was down can be recovered from
- the order, as long as something goes looking. */ export function reconcile(fetched: readonly ShopifyOrder[], recorded: readonly Recorded[]): Finding[] { // Retries and duplicate deliveries mean one order can be recorded several // times; what matters is whether it was recorded at all. const seen = new Set(recorded.map((r) => r.orderId));
const findings: Finding[] = [];
for (const order of [...fetched].sort((a, b) => a.createdAt.localeCompare(b.createdAt))) {
if (seen.has(order.id)) continue;
const ref = order.sellerRef?.trim();
findings.push(ref ? { kind: "missed", orderId: order.id, sellerRef: ref } : { kind: "unattributed", orderId: order.id });
}
return findings;
}`
This only works because of a decision made much earlier: the seller's identity travels on the order itself, as an attribute, rather than living only in the app's memory of a webhook. That is what makes an outage recoverable. If attribution existed only in what the app was told, four hours of downtime would be four hours of sales nobody could ever credit; because it is on the order, the answer is still sitting in Shopify, waiting to be fetched. The rest is bookkeeping — deduplicate by order rather than by delivery, and report orders with no seller instead of dropping them, since a silent skip is how a missing attribute becomes a missing payout.
Make the windows overlap. A daily job that looks back exactly twenty-four hours has a seam at midnight, and an order placed while the previous run was still working can fall into it. Looking back forty-eight hours every day costs almost nothing, because the comparison is safe to repeat — an order already recorded is simply skipped, however many times the job sees it.
Overlap is only cheap when the processing is idempotent, which is one more reason to build it that way.
The same job is also the cheapest monitor you'll ever write. A reconciliation that finds gaps every day is telling you the webhooks are failing, long before anyone asks why their total looks low. It should check the subscriptions still exist too, because after a long enough outage they may not.
The defaults are built for getting started
Worth knowing what you're inheriting. Shopify's app template stores sessions in SQLite through Prisma, and its own deployment guide notes you can only run more than one web container if the database gets its own container or volume. The same guide warns that one popular host may suspend idle containers and reset disk storage — fine for a demo, a quiet catastrophe for a database living on that disk.
Managed hosting isn't automatically safer. It moves the failure somewhere you're less likely to look.
The address is part of the contract too. When the app's URL changes, the configuration has to be updated and redeployed with shopify app deploy. Moving a self-hosted app to a new server isn't only a server migration — until Shopify is told, the admin is embedding the old address and every webhook is going to it.
What self-hosting doesn't cover
The server. Patching, firewalls, TLS renewal and restore drills are yours now, and a backup nobody has ever restored is a hope rather than a backup.
Data obligations. Shopify makes its privacy compliance webhooks mandatory for App Store apps, with thirty days to act. A custom app on one store isn't named there, but the personal data it holds is still personal data.
Reconciliation, on any host. A managed platform reduces how often the app is down. It does not make it never down.
When this is the right shape
It isn't, if an App Store app does the job. Someone else hosts it, monitors it, and answers for it — for most needs that's a better deal than any server you could run.
It's right when the app holds records the business couldn't rebuild from Shopify alone — money, attribution, an append-only ledger — and the business wants to own where those records live. Then hosting, backups, monitoring and reconciliation are part of the build, not chores after it.
The cost should be said out loud at the start: you now own uptime. Nobody at Shopify will notice the app is down. Nobody at a hosting company will restart it. If the process stops, it stays stopped until a human or a monitor notices.
Takeaways
Shopify hosts the store, not your app. A custom app needs HTTPS, credentials, a database and a process that stays up — and nobody at Shopify will notice if it doesn't.
A webhook gets five seconds and up to eight retries over four hours. If failures persist, the subscription is removed.
Webhooks can arrive more than once. Deduplicate by the event, not the delivery.
Webhooks are the fast path, not the record. Reconcile against the orders on a schedule, which only works if what you need is stored on the order.
Self-hosting suits an app holding records the business cannot rebuild elsewhere. The price is owning uptime, patching and restore drills.
Two things I'd like other people's experience on.
The five-second budget, in practice. Acknowledge-then-work is the obvious answer, but it means the webhook handler returns 200 before anything is durable — so a crash between the acknowledgement and the write loses the event, and now reconciliation is the only thing that catches it. The alternative is writing to a queue before acknowledging, which is durable but puts a write back inside the budget. I've gone with the queue. Interested in whether people find the bare acknowledge-then-work version holds up under real load.
Reconciliation windows. Forty-eight hours covers the midnight seam cheaply. It doesn't cover a weekend outage, which is exactly when nobody is watching. I've not found a window that's both cheap enough to run daily and wide enough to catch the outages that actually happen — so the real answer is probably monitoring rather than a bigger window, and the reconciliation job is just the backstop.
If you run a self-hosted Shopify app, how long did it take before the first outage taught you something the design hadn't anticipated?
Top comments (0)