DEV Community

orbiresearch
orbiresearch

Posted on Originally published at orbiresearch.com

Approval timeouts: what your agent should do when nobody answers

§ 01 · The missing fifth field

In the approval queue pattern I described four fields every approval item should carry: the action, the justification, the blast radius, and the confidence with an alternative. A reader on dev.to pointed out what was missing, and they were right. None of those fields answer a simple question: what happens when nobody answers?

Every real queue has items that age. People are in meetings, on holiday, or just busy. If you don't decide up front what an expired item turns into, the system decides for you, and it usually decides badly.

§ 02 · How queues rot without a timeout policy

There are two common ways this goes wrong.

◆ The graveyard. Nothing happens on timeout. Items pile up, the agent stalls, and after a few weeks people learn that the queue is where requests go to die. They stop looking at it, which is the opposite of what you built it for.

◆ The global auto-approve. Someone gets tired of the backlog and adds a rule: anything older than 24 hours gets approved. It feels practical. It also means the most dangerous item in the queue, the one nobody felt comfortable approving, now goes through automatically. You have rebuilt the rubber stamp, just with a delay.

In the original post I wrote that an SLA breach should escalate to a person and never auto-approve. I'd now make that more precise. For irreversible actions, that still holds. For some action types, a different default is the better answer.

── Defaults by reversibility ──

§ 03 · One default per action type

The fix is to give each action type its own timeout and its own default action, and to choose that default by asking one question: how expensive is it to undo this if the default turns out to be wrong?

In practice four defaults cover almost everything.

◆ Approve. For actions that are internal, reversible and low impact. Posting a summary to an internal dashboard, tagging a ticket, saving a draft. If nobody objects in time, the cost of doing it is close to zero and easy to reverse.

◆ Hold and escalate. For actions that leave your system or move money. Sending an email to a customer, issuing a refund, paying an invoice. These never go through on silence. On timeout, the item goes to the next person on the escalation list and keeps waiting.

◆ Cancel. For actions whose value expires. Booking a slot for today, replying to a quote request after the customer's deadline, reacting to a price that has already changed. Doing them late can be worse than not doing them. On timeout, the item is closed and the requester is told.

◆ Re-plan. For actions built on data that may now be stale. If a proposal was based on a stock level, a balance or a status from yesterday, approving it today can be wrong even if it was right when it was created. On timeout, the item goes back to the agent, which checks the facts again and creates a fresh request if it still makes sense.

The human who owns the queue picks the default per action type, in writing, before anything goes live. It is a product decision, not an engineering detail.

── Picking the numbers ──

§ 04 · Timeouts are per type, not global

One global SLA does not work, because a refund and a dashboard update do not deserve the same patience. Start with rough values per action type, for example 30 minutes for time sensitive customer replies, 4 working hours for refunds and 24 hours for internal changes, then tune them from real data.

The useful numbers are the median and the 90th percentile of how long approvals actually take for each type. If the 90th percentile is longer than the timeout, the timeout will fire on normal days, and people will start to ignore it. If the median is only a few seconds, reviewers are probably not reading, which is a different problem and a more serious one.

§ 05 · A small implementation

The policy is small enough to live next to the routing rules. A minimal version:

from dataclasses import dataclass
from datetime import datetime, timedelta
from enum import Enum

class OnTimeout(Enum):
    APPROVE = "approve"
    ESCALATE = "escalate"
    CANCEL = "cancel"
    REPLAN = "replan"

@dataclass(frozen=True)
class TimeoutPolicy:
    after: timedelta
    on_timeout: OnTimeout
    escalate_to: str | None = None

POLICIES = {
    "internal_dashboard_post": TimeoutPolicy(timedelta(hours=24), OnTimeout.APPROVE),
    "customer_email": TimeoutPolicy(timedelta(minutes=30), OnTimeout.ESCALATE, "support_lead"),
    "refund": TimeoutPolicy(timedelta(hours=4), OnTimeout.ESCALATE, "finance_lead"),
    "same_day_booking": TimeoutPolicy(timedelta(hours=2), OnTimeout.CANCEL),
    "stock_reorder": TimeoutPolicy(timedelta(hours=12), OnTimeout.REPLAN),
}

def resolve_expired(item, now: datetime):
    policy = POLICIES[item.action_type]
    if now - item.created_at < policy.after:
        return None  # still waiting
    return policy.on_timeout, policy.escalate_to
Enter fullscreen mode Exit fullscreen mode

Two details matter more than the code. Any action type that is not in the table should fall back to ESCALATE, never to APPROVE. And every decision made by a timeout has to be logged as a timeout decision, with the policy that caused it, so that an audit can tell the difference between "a person approved this" and "nobody looked and the clock approved it".

── What you get ──

§ 06 · What changes when you do this

The queue stays short. Low stakes items drain on their own, expired time sensitive items close cleanly, and stale ones get rechecked instead of approved blind. The only items that actually age are the irreversible ones, which are exactly the ones you want a person to look at.

That is also why reviewers keep reading. When the queue only holds decisions that matter, people treat it with the attention it needs.

If you want to try this, add a fifth column to the table in your APPROVALS.md: the timeout and the default for each action type. It takes an hour to write and saves a lot of arguments later.

Thanks to the reader who raised this. The reference implementation of the other patterns is on GitHub (agent-reliability-patterns), and the original post is here (The approval queue pattern).

── End of pattern ──

◆ Every action type needs a timeout and a default: approve, hold and escalate, cancel, or re-plan.

◆ Choose the default by how expensive the action is to undo. Anything unknown escalates.

Log timeout decisions separately from human ones. "Nobody objected" is not the same as "someone approved".

ORBIRESEARCH


Originally published on the OrbiResearch Lab. We build production AI agents at orbiresearch.com.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.