DEV Community

life
life

Posted on

A Cancel Is a State Transition, Not a Boolean

Almost every fulfillment integration I have seen handles a cancellation by flipping a flag on the order row and hoping the warehouse worker process polls often enough. It works in demos and in quiet months. It fails in the specific way that costs money. The flag says canceled, the parcel leaves anyway, and both sides of the integration are technically correct about what they did.

The fix is not a faster poll. It is admitting that a cancel is an intent competing with a state machine, and that the state machine has to be the thing that resolves the race.

The two tables

The order record already has a status, and Shopify's own vocabulary for it is open, closed, canceled and archived, with a separate fulfillment axis of unfulfilled, partially fulfilled and fulfilled. Copying those into your database is fine. Treating them as one field is the bug, because they change for different reasons and at different speeds.

What is missing is the intent record:

create table cancel_intent (
  id              bigserial primary key,
  order_ref       text not null,
  source          text not null,          -- storefront | support | fraud | customer
  reason_code     text not null,
  received_at     timestamptz not null default now(),
  observed_state  text,                   -- what the resolver found
  outcome         text not null default 'pending',
                                    -- stopped | too_late | unwind_required | failed
  resolved_at     timestamptz
);
Enter fullscreen mode Exit fullscreen mode

The column that makes the whole design honest is observed_state. It forces the system to record what was true at the moment the cancel arrived, instead of pretending that a cancel is a command. A cancel is a request to move a parcel backwards through states that may already be behind you.

Resolving the race

The resolver runs once per intent, takes a lock on the fulfillment row, and branches on the current state. The important detail is that the pick confirmation and the cancel have to serialize on the same row, otherwise you get the classic interleaving where the picker reads "not canceled", the cancel writes, the picker writes "picked", and the parcel is on a truck with a canceled order attached to it.

lock fulfillment_row
switch row.state:
  case 'queued':        state = 'canceled'; outcome = 'stopped'
  case 'picked':        emit walk_back_task;  state = 'canceled';
                        outcome = 'unwind_required'
  case 'packed':        void_label; emit unpack_task;
                        state = 'canceled'; outcome = 'unwind_required'
  case 'manifested':    outcome = 'too_late';  -> return_path
  case 'shipped':       outcome = 'too_late';  -> return_path
Enter fullscreen mode Exit fullscreen mode

Two of these branches are not cancellations at all. They are unwind operations that happen to have been triggered by a cancellation, and they need their own tasks, their own completion signal, and their own effect on the stock record. Modeling them as the same event is how you end up with a canceled order and an on-hand quantity that never came back.

The last two branches are where the honest system says no. Once a carton is on a manifest, the parcel belongs to a carrier schedule, and the remaining option is whatever intercept product that carrier sells, which is time-boxed and explicitly not guaranteed. Your state machine should not pretend otherwise, because the next person to read that record will be reconciling money.

The acknowledgment is the product

Support tickets die here. The storefront gets 200 OK from a cancel endpoint and reports "canceled" to the customer, while the warehouse got the message at 14:38 and the carton was sealed at 14:36. Both are true.

An ack that carries the observed state turns that argument into a log line. "Cancel received at pack-out. Label voided, unit walked back to the shelf, count reconciled." That is a sentence a support agent can send, and it cannot be written unless the resolver actually ran. The difference between those two replies is the entire service level.

The same event should carry a reason code. Cancels arrive from customers, from fraud review, from support goodwill and from marketplace rules, and they have different follow-up obligations. A cancel that came from a marketplace deadline has to be visible to whoever watches channel compliance, while a customer change of mind does not.

What this fixes that a flag does not

Three things show up in the data afterwards. Late notices stop being invisible, because too_late is a countable outcome with a stage attached. The unwind tasks stop being forgotten, because they have a row that is either complete or not. And the cost of a cancel becomes measurable per stage rather than per month, which is the only version of that number that changes anybody's behavior.

We run this shape across our Chinese sites, 13,000 square meters in Suzhou, 3,000 in Shenzhen and 8,000 in Dongguan, under the FulfillNexa by SBT (fulfillnexa.com) name, and the reason it is modeled as transitions rather than flags is not architectural taste. It is that a cancel can arrive while a human is holding the unit, and the record has to say what happened next in a language that survives the shift change.

The status names and the rule that a partially fulfilled order has to have its fulfillment canceled before the order itself becomes eligible are Shopify's published documentation. Check their current version the day you build, because the field that saves you is the one that still exists.

Top comments (0)