In my last post, I walked through one problem where durable execution fit naturally. This time I want to zoom out. Some of the best candidates for this pattern aren't exotic. They're sitting in products we use every day, and most teams build them the traditional way without a second thought.
One note: these are my own sketches of how such systems could be modeled, not how any of these companies actually build them. The code is pseudocode, not production code.
The Pattern We Keep Missing
When we build a feature, we usually follow the flow of what the system already has: a Kafka topic here, a status column there, a cron job somewhere else. Over time, code for orchestration and business code get fused together. Retries, acknowledgements, timers and state sync end up tangled with the actual product rules.
A fresh look at these use cases lets us spot the workflow shape early, in the initial version of a product. Then the separation isn't only about infrastructure. Whole behavioral patterns can be pulled out of the backend and treated as one unit.
Use Case 1: The Booking That Lives for Weeks (Airbnb-style)
A traveler confirms a booking, and from that moment a reservation state machine starts:
- Payment: the account is debited on a set day after confirmation.
- Refund window: fully refundable until 48 hours before check-in. After that, cancellation is still possible, but only for the remaining nights.
- Policy varies per listing, so the rules aren't global.
- Human in the loop: a cancellation on the day of arrival needs a decision from the host or traveler. We want humans involved only for that decision, nothing else.
async function reservationWorkflow(booking, policy) {
let cancel = null;
onSignal('cancelRequested', (req) => { cancel = req; });
await sleepUntil(policy.debitDate);
await chargeGuest(booking);
await sleepUntil(booking.checkIn - hours(48)); // free-cancel window closes
// ...partial refund rules apply from here
if (cancel && isArrivalDay(cancel)) {
const decision = await waitForSignal('hostDecision'); // human step
await applyResolution(booking, decision);
}
}
The timeline of the booking is the code. Humans only show up where judgment is needed.
Use Case 2: The Order and Its Journey (E-commerce)
This one is everywhere. A delivery partner like DHL sends webhooks about a package's status. The usual approach:
- The webhook lands on a Kafka topic.
- One service picks it up, processes it, and produces a follow-up event.
- Another service picks that up, decides whether this user even wants this notification ("out for delivery" yes, "arrived at intermediate hub" probably not), and sends it.
- Retries, acknowledgement checks and dedupe are scattered across all of it.
Business rules and orchestration end up in the same handlers. Changing "who gets which notification" means touching retry logic too.
Now flip the view. The order is an entity, and there's one long-running workflow on it, from the seller's warehouse to the customer's door:
async function orderJourneyWorkflow(orderId, prefs) {
let state = await initialState(orderId);
onSignal('carrierUpdate', async (update) => {
state = applyUpdate(state, update);
if (prefs.wants(update.level)) {
await notifyUser(orderId, update);
}
});
await waitUntil(() => state.delivered || state.cancelled);
}
All the handlers live in one place, and they read like plain API calls instead of "Kafka subscribe, Kafka publish, blah blah."
looks like a React Component?
This is the mental model that clicked for me. A React component lives for maybe two minutes while a user interacts with it. It holds state, has handlers for events, and each handler says how state changes.
A durable workflow is the same shape, but the "session" stretches over days or weeks, and the "clicks" come from carriers, hosts, timers or other services. The business logic is only this: given the current state and this interaction, what's the new state? Persisting state, queuing events and waking up at the right time are abstracted away.
Bonus: Shared Work, Done Once
Not every workflow is per-user or per-order. Take a streaming home screen: "Top 10 in your region" depends on the segment a user belongs to, not the user. There are far fewer segments than users, so for each segment and each interval you want to compute once and serve thousands. A small workflow per segment (gather, rank, publish, sleep until the next interval) keeps that calculation from running thousands of times, with durable state if it fails halfway. Fair warning: the heavy compute lives elsewhere, so the win here is orchestration, not the ranking itself.
Does It Have to Be Temporal?
No. You could build this yourself with a persistence layer for state and a queue for events. Temporal is one way, and there are others emerging. The tool matters less than the realization: software keeps getting more complex, and we keep working on smaller pieces of bigger systems. That's a good moment to pull complete packages of functionality into one place where they're maintainable, and to solve their shared problem (orchestration) once, together. Consistency is a bonus.
Cheers!!
Top comments (0)