Disclosure: I work on 24hTrack, a free multi-carrier package tracker. The numbers below come from parcels tracked on our platform. The modelling problem applies whatever tracker you use.
Every post-purchase system eventually grows a stale-shipment alert. The logic is always the same shape:
if (now - lastScanAt) > SEVEN_DAYS -> flag as stalled
It is the first rule anyone writes, and on cross-border volume it is close to useless. Here is the measurement that killed it for us.
The measurement
For every parcel that was eventually delivered, take the single longest stretch between two consecutive carrier scans. Not the average gap — the worst one, the number your alert is actually racing against.
Across a 120-day window:
| Longest silent gap reached | Share of delivered parcels |
|---|---|
| 3 days or more | 58% |
| 5 days or more | 44% |
| 7 days or more | 27% |
| 10 days or more | 11% |
| 14 days or more | 4% |
Median 4.1 days. p90 10.3 days. p95 13.1 days.
So a seven-day threshold fires on more than a quarter of parcels that arrive perfectly fine. If a quarter of your alerts are noise, nobody reads the queue, and the 4% that are genuinely stuck get the same treatment as the 27% that are not.
Why the gaps exist at all
A scan is a physical event: something passes a reader. Five stages of a normal journey have no reader in them.
- Waiting in an export warehouse for space on a flight
- The flight
- Import customs — a queue, not a conveyor
- The handover between two carriers
- The final van run, where the next scan is the doorstep
None of these is an anomaly. They are the route. A model that treats elapsed-time-since-last-event as a health signal is measuring scan density, not parcel health — and scan density is a property of the carrier's network design.
That shows up directly in the per-carrier numbers
Same measurement, split by carrier (median longest gap on delivered parcels, and the p90):
| Carrier | Median | p90 |
|---|---|---|
| Rayspeed Asia | 1.0 d | 2.6 d |
| FedEx | 2.1 d | 19.4 d |
| Cainiao | 2.7 d | 9.9 d |
| UPS | 3.0 d | 10.0 d |
| Yanwen Express | 3.2 d | 10.1 d |
| China Post | 3.9 d | 9.5 d |
| USPS | 4.6 d | 10.2 d |
| UniUni | 6.3 d | 12.5 d |
| CTT Express | 11.0 d | 17.8 d |
Look at FedEx: median 2.1 days but p90 19.4. A single global threshold cannot serve a distribution that skewed. Set it at 7 and you spam; set it at 20 and you miss two and a half weeks of a genuinely dead parcel on a dense-scanning lane.
Full dataset, CC BY 4.0: silent-gap-by-carrier.csv
What to build instead
1. Threshold per carrier, from your own p90 — not a constant.
// carrierGapP90: { 'USPS': 10.2, 'UniUni': 12.5, ... }, recomputed monthly
const stalled = daysSince(lastScanAt) > (carrierGapP90[carrier] ?? GLOBAL_FALLBACK)
Recompute from delivered parcels only. Parcels still in flight censor the distribution downwards; parcels that died censor it upwards. Only completed journeys give you an honest denominator.
2. Condition on the leg, not just the clock. The same 10 days means different things before and after the parcel reaches the destination country. A domestic last leg that goes quiet for 10 days is a real signal; 10 days between an export scan and an import scan is a flight plus a customs queue.
3. Promote status text over elapsed time. These beat any gap threshold, and none of them is time-based:
-
customs+documents required/ duties owed — nothing moves until the buyer acts, and nobody chases them -
return to sender,refused,undeliverable - a delivery attempt with no follow-up scan
- no first scan at all — a label exists but nothing was ever handed over. This is the cheapest high-precision signal in the whole system and most implementations skip it, because the parcel has an event and therefore looks tracked.
4. Alert on transitions, not on states. A parcel in a quiet stretch is in a state. A parcel that just acquired an exception scan made a transition. Only the second one is news.
The downstream reason this matters
Support macros inherit the alert's thresholds. If your system calls a parcel stalled at day 7, your agent tells the customer something is wrong at day 7 — on a parcel with a 27% chance of being completely normal. You have now created the ticket you were trying to prevent, and the follow-up ticket when it arrives anyway.
The reply that actually deflects contains a number: "no new scan yet, and that is normal up to about X days on this route; here is what a real problem looks like." You can only write that sentence if you measured X.
24hTrack (24htrack.com) is a free package tracker for 3,200+ carriers: paste any tracking number, the carrier is detected automatically, no sign-up. There is a REST API and an MCP server if you want this in your own stack, and the open guide plus transit-time data are on GitHub.
Top comments (0)