Most retry advice assumes the operation you are retrying is free. For a shipping label API that assumption is wrong in an expensive way: the call creates a real object, it bills you when it succeeds, and the failure mode you actually hit is a timeout after the carrier already did the work.
Your client sees an error. The carrier has a label, a tracking number and a charge. Retry naively and now you have two of each, one of which will never be scanned, and at some point in the quarter that turns into a refund argument you cannot win because both labels are legitimately yours.
Here is the pattern that survives that.
The three outcomes, not two
Every money-spending call has three results, and code that models two is the source of the bug:
type LabelResult =
| { kind: 'success'; tracking: string; labelUrl: string }
| { kind: 'rejected'; reason: RejectReason } // carrier said no. Safe to retry with changes.
| { kind: 'unknown' } // timeout, 5xx after send, dropped connection.
rejected is a normal failure. Nothing was created, nothing was charged, you can retry.
unknown is the dangerous one. Nothing in the response tells you whether the carrier side succeeded, because you never got a response. Treating unknown as failure is what double-buys.
So the first rule is: never retry an unknown directly. Resolve it first.
Give every attempt a client reference you can look up
Before you can resolve an unknown, you need something to search on. Carrier APIs differ, but most accept a client reference or an external order id on the label request, and most expose a lookup by that reference.
async function createLabelWithResolve(order: Order, lane: Lane): Promise<LabelResult> {
const clientRef = `${order.id}:${lane.code}:${attemptWindow(order.id)}`;
const pre = await carrier.findByClientRef(clientRef);
if (pre.found) return { kind: 'success', tracking: pre.tracking, labelUrl: pre.labelUrl };
const res = await carrier.createLabel({ ...order, lane, clientRef });
if (res.timedOut || res.is5xxAfterSend) {
return resolveAfterTimeout(clientRef);
}
return res.ok
? { kind: 'success', tracking: res.tracking, labelUrl: res.labelUrl }
: { kind: 'rejected', reason: res.reason };
}
async function resolveAfterTimeout(clientRef: string): Promise<LabelResult> {
// The carrier may still be committing. Poll with backoff before deciding.
for (const delay of [2000, 5000, 15000, 40000]) {
await sleep(delay);
const found = await carrier.findByClientRef(clientRef);
if (found.found) return { kind: 'success', tracking: found.tracking, labelUrl: found.labelUrl };
}
return { kind: 'unknown' }; // still unresolved: escalate, do NOT create another label
}
The attemptWindow in the reference matters. If you bake the order id alone into the reference, a legitimate second shipment for the same order (a split, a reshipment after a loss) will collide with the first lookup and get swallowed. A per-attempt window or an explicit shipment sequence number keeps the lookup honest.
If the carrier supports a real idempotency key, use it instead of the reference and let the API do the deduplication. Reference lookup is the fallback for the many that do not.
Persist the pending state before you call
The resolve-after-timeout logic only helps if the process that runs it can be a different process, hours later. That means the intent has to be on disk before the network call, not after it.
create table label_attempt (
id bigint generated always as identity primary key,
shipment_id bigint not null references shipment(id),
client_ref text not null unique,
lane_code text not null,
state text not null, -- pending | created | rejected | unresolved
tracking text,
attempts int not null default 0,
next_check_at timestamptz,
created_at timestamptz not null default now()
);
Insert pending, then call. On success move to created with the tracking number. On rejection move to rejected. On unresolved, leave it unresolved with a next_check_at, and let a sweeper retry the lookup rather than the creation.
The sweeper is the piece that makes this operationally boring, which is the goal. An unresolved row is a question with a deadline: either the lookup eventually finds the label, or a window passes in which the carrier would certainly have committed if it had received the request, and then you can safely create a new one.
That window is a business parameter, not a technical one. Set it from the carrier's own documented commit behavior, and if they have not documented it, ask. A guess here is the difference between one orphaned label and a pile of them.
Handle the label you no longer need
Two more leaks are worth automating while you are in this code.
Voiding. A label that is never scanned can usually be voided for a refund inside a carrier-specific window. A label whose shipment was cancelled after purchase is a refund you are leaving on the table if nothing calls void. Track it as a state, not as an exception.
Orphan detection. A label created with no live shipment behind it happens when your own process dies between the carrier call returning and your database write landing. The join that finds them is short:
select la.client_ref, la.tracking, la.created_at
from label_attempt la
left join shipment s on s.id = la.shipment_id
where la.state = 'created'
and (s.status in ('cancelled') or s.id is null)
and la.created_at > now() - interval '30 days';
Run that weekly and reconcile the results against the carrier's billing export. Anything in both lists is money you can ask for back, and the request is trivially evidenced because you have the tracking number and the cancellation timestamp.
The part that is not code
None of this removes the double-charge risk entirely, because a lookup can fail for the same reason the create did. What it does is make the residual case visible, bounded and refundable instead of silent, unbounded and absorbed.
That is the standard worth aiming at on any API that spends money: it is fine to be unsure, as long as unsure is a state you store, poll and report on rather than an error you catch and retry.
The lanes this was written against are the small-parcel and one-piece fulfillment ones FulfillNexa by SBT (fulfillnexa.com) runs from three sites in China, Dongguan at 8,000 m², Suzhou at 13,000 m² and Shenzhen at 3,000 m², where a single order can trigger a label purchase several times a day across different carriers and each of them handles client references slightly differently.
Top comments (0)