Most tracking integrations are written as if a shipment has one carrier. Register the number, poll or receive a webhook, map the status, show a timeline. That model is fine for a domestic parcel and quietly wrong for a cross-border one, and the place it breaks is the handover — the moment the origin carrier stops holding the box and a courier in the destination country starts.
I work on 24hTrack, a free multi-carrier tracker covering 3,200+ carriers where you paste any tracking number and the carrier is detected automatically. What follows are measurements from delivered parcels on that platform over the 120 days to 2026-09-24, published as CC BY 4.0 in this dataset. They changed three assumptions I had built into tracking code.
1. A cross-border number keeps reporting after the handover — usually
The intuitive model is that the export carrier's number goes dark at the border and a second number takes over. Sometimes. But look at scans published per shipment:
| Shipment type | Median scans |
|---|---|
| 4PX (cross-border) | 20 |
| Yanwen Express (cross-border) | 19 |
| China Post · Cainiao | 18 |
| FXYL (cross-border) | 17 |
| USPS (domestic) | 7 |
A cross-border number publishes two to three times as many events as a domestic one. Those extra events are largely the destination side of the journey, written back under the original number. So:
Do not treat "no second tracking number" as an incomplete record. In most networks there is nothing to link, because the one number already spans both legs. A data model that requires a child_tracking_number to consider a shipment complete will mark the majority of healthy cross-border shipments as broken.
2. The handover gap will fire your stale-shipment alert
The usual staleness rule is "no event in N days → alert". The handover is exactly where that misfires, because it is a custody change with no scanner in it: the origin carrier has stopped scanning and the destination carrier has not started.
Across delivered parcels, the single longest silent stretch per parcel runs to a median of about 4 days, and more than 1 in 4 went a full week with no scan before arriving fine. A flat 7-day rule fires on roughly a quarter of successful shipments.
Two fixes that cost very little:
# instead of
if now - last_event > 7d: alert()
# use a per-carrier threshold from that carrier's own p90
if now - last_event > carrier_p90_silent_gap: alert()
# and suppress entirely while the last event is a known handover/customs state
if last_event.category in {CUSTOMS, HANDOVER, ARRIVED_DESTINATION_COUNTRY}: suppress()
The p90 varies enormously — about 2.6 days for one operator, 19.4 for another — so a single global constant is the worst possible choice.
3. The final transit scan is a much stronger ETA signal than elapsed time
This is the one I would actually change first. Measure the gap between the last transit scan and the delivery scan:
| Carrier | Median | 80% within | 90% within |
|---|---|---|---|
| USPS | ~5 h | ~8 h | ~10 h |
| Yanwen Express | ~8 h | ~12 h | ~34 h |
| China Post | ~8 h | ~46 h | ~84 h |
| Cainiao | ~8 h | ~31 h | ~91 h |
| 4PX | ~8 h | ~17 h | ~58 h |
| FXYL | ~1 d | ~26 h | ~29 h |
| YunExpress | ~26 h | ~79 h | ~96 h |
Nearly all of the variance in a cross-border delivery lives before that scan — in flight consolidation, import customs and sorting backlogs. After it, the distribution collapses to hours.
That makes "has the final transit scan landed?" a far better feature than "days since dispatch" for a delivery-window prediction, and a much better trigger for a "arriving today/tomorrow" notification. Elapsed time since dispatch is mostly noise from customs queues you cannot observe.
4. Two clocks, not one, when you show an estimate
A concrete pair from the same data. An export-side operator (FXYL, a Chinese air-freight forwarder) runs a median 22.8 days from first carrier scan to delivery, 80% within 26.8. A last-mile carrier in the United States (YWE) delivers half of its parcels within about 2 days of its own first scan, 80% within 5.
If your UI shows one progress bar for "shipping", you are averaging those two distributions together and will be wrong at both ends. Modelling the legs separately — international leg, then domestic leg keyed off the arrival scan — gets you an estimate that does not embarrass you on day 20.
What I would build differently
- Store the raw carrier status string, not just your normalised bucket. You cannot re-derive "handed to local partner" from
IN_TRANSITlater. - Make staleness thresholds per-carrier and data-derived, refreshed periodically rather than hard-coded.
- Key the "almost there" signal off the last transit scan, not off elapsed days.
- Do not model a second tracking number as required. Support it when it exists, never depend on it.
- Expect an event ordering you did not choose: some carriers publish newest-first, some oldest-first, and a few backfill history when a new leg starts.
Dataset and the full write-up are public: final-leg-by-carrier.csv and the guide, both CC BY 4.0. If you want to sanity-check a number against a live multi-carrier lookup while building, 24hTrack does it free without an account, and has a REST API and MCP server if you want it inside an agent.
Top comments (0)