DEV Community

Agent Island Pro
Agent Island Pro

Posted on Originally published at agentislandapp.github.io

How to know when Claude Code is done

You start a long task, switch to something else, and come back nine minutes later to find it finished eight minutes ago — or worse, that it stopped to ask you a question and has been sitting there since.

There is a built-in notification for this. It is off in most terminals, and it only fires when Claude Code believes you are away.

Settings below are from Claude Code's own documentation, September 2026.

The part that is already built in

Claude Code fires a notification when it finishes a task or pauses for a permission prompt — with a condition that explains most of the "it never notifies me" reports: it only fires when you appear to be away from the terminal.

If you are sitting in the window watching it work, there is nothing to interrupt you about, so nothing happens. Testing it by staring at the terminal is therefore a test designed to fail.

The second condition is the terminal itself. A desktop notification is sent by default in Ghostty, Kitty and iTerm2, and nowhere else. In any other terminal — Warp, the VS Code integrated terminal, Apple Terminal, Alacritty — nothing arrives until you ask for the bell instead:

// ~/.claude/settings.json
{
  "preferredNotifChannel": "terminal_bell"
}
Enter fullscreen mode Exit fullscreen mode

Three further things swallow the notification even once it is enabled, and each looks like the feature being broken:

If Then
You use iTerm2 Forwarding is not on by default. Settings → Profiles → Terminal, enable Notification Center alerts, then under Filter Alerts allow escape-sequence-generated ones.
You run inside tmux tmux swallows it. set -g allow-passthrough on in ~/.tmux.conf, then tmux source-file ~/.tmux.conf.
Nothing appears anywhere Check the terminal application itself has notification permission in your OS settings — the notification is delivered by the terminal, not by Claude Code.

If you would rather have a sound you choose, a Notification hook runs alongside the built-in one rather than replacing it:

{
  "hooks": {
    "Notification": [
      { "hooks": [{ "type": "command",
                    "command": "afplay /System/Library/Sounds/Glass.aiff" }] }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

That is the whole of the built-in answer, and for a single session it is usually enough. The rest of this is about why it stops being enough, and it starts with a definition problem.

"Done" is three states, not one

If you build anything on top of this — a notifier, a status line, a dashboard — the first thing you discover is that "is it done?" has no single answer. There are three states, they mean different things to you, and only one is urgent:

State What it means Does it need you?
Running Working. Tools are being called. No
Waiting Stopped mid-task on a permission prompt, a question, or a plan to approve. Yes, and it is blocked until you answer
Idle The turn ended. It said its piece and is waiting for your next instruction. Only when you want it to

The events map onto those states cleanly enough once you know which is which:

  • UserPromptSubmit and PostToolUse → running. Each is also proof of life, which matters later.
  • Notification → waiting. The message says whether it is a permission request or a question.
  • Stop → idle. Not "the work is finished" — it means this turn ended, which is a much smaller claim.
  • SubagentStop → nothing, as far as the parent is concerned. A subagent finishing says nothing about whether the session that spawned it needs you.

The distinction that matters most is Waiting versus Idle, because they feel identical from the outside — in both cases the terminal has gone quiet — and they could not be more different. Idle means you can look whenever you like. Waiting means it stopped mid-task and nothing will happen until you answer, which is the case where a nine-minute delay is nine minutes wasted.

A notifier that treats "quiet" as one state will either nag you about every finished turn or let you sit on a blocked prompt.

The trap: a status nobody is going to correct

Here is the failure I shipped, and it is the one worth stealing the fix for.

On 21 August a session sat showing waiting for your input for over an hour. The user had answered in the terminal within seconds. Nothing was wrong with the answer — the problem was that the hooks had been connected after that session started, so the event that would have cleared the waiting state was never sent. The state was correct when it was written and quietly became a lie afterwards.

This is not a niche case. Any of these produce the same shape:

  • The session started before your hooks were installed
  • The process was killed, so no Stop and no SessionEnd ever arrives
  • The machine slept mid-task
  • Your hook script broke for ten minutes and you fixed it

An event-driven state machine only knows what it was told. It cannot distinguish "still waiting" from "was waiting, and the update went missing", because both look like silence. So it needs one rule that does not depend on receiving anything:

// After this long with no event at all, a non-idle status
// is no longer evidence of anything.
static let staleAfter: TimeInterval = 10 * 60

func isStale(now: Date) -> Bool {
    status != .idle && now.timeIntervalSince(lastEventAt) >= staleAfter
}
Enter fullscreen mode Exit fullscreen mode

Ten minutes of total silence and we stop asserting the state. Note the status != .idle half: idle is the one state that should persist, because a session that ended its turn an hour ago is still, correctly, a session that ended its turn. Only the active claims expire.

The principle generalises past this codebase: a status you cannot stand behind must not be presented as current, and stale evidence should decay on its own rather than wait for a correction that may never come.

Two hooks, opposite designs

One last thing worth copying if you write your own. Status events and permission prompts want opposite designs, and using one design for both is how you end up with a notifier that makes your agent feel slow.

Status events Permission prompt
Job Report what happened Answer a question, or decline to
Design Fire and forget Blocking
My timeout curl -m 0.3 curl -m 52, inside a configured "timeout": 55
On failure Give up silently, exit 0 Print {} — no opinion — and exit 0 (exit 2 would block)

A status hook must never make the CLI wait: nothing depends on its answer, so a third of a second is generous and anything longer is a tax on every single tool call.

A permission hook is the opposite — something is waiting on its answer — but it still has to return inside whatever timeout it was given. The default for a command hook is 600 seconds, which is much longer than you want a blocking hook to hold a tool call, so set it deliberately.

Both exit 0 on an ordinary failure, and deliberately so, because exit 2 would block the call. There is more on that in why your Claude Code hook isn't running.


Originally published at agentislandapp.github.io/done.html. I write these while building AgentIsland, a macOS app that puts this state in the notch: a dot per session while things are running, and the actual permission prompt the moment one is waiting.

Top comments (8)

Collapse
 
hamid_ahmadian_3570449f72 profile image
Hamid Ahmadian •

The Running/Waiting/Idle split maps cleanly right up until you're running more than one session at once, which is the case a notifier like this eventually has to handle — a worktree per feature branch, or a few tabs open against the same repo. Every hook payload carries a session ID (and cwd), but nothing in the built-in notification path surfaces which pane it came from, so a naive implementation that just listens for "Notification fired, go Waiting" collapses five sessions into one status the moment two of them are active at once. The fix ends up being the same shape as your stale-state rule: key the state machine by session ID, not by "the terminal," and treat a session you haven't heard from in your staleAfter window as gone rather than merging it into whatever session did just report. Worth calling out explicitly for anyone extending this past a single-session dashboard, because the failure mode looks identical to the stale-state bug in the post — a status that was true for session A gets silently read as true for session B — but the fix is routing, not aging.

Collapse
 
agentislandpro profile image
Agent Island Pro •

Exactly the failure that matters, and the (tool, session_id) framing is the part I'd underline: session_id alone isn't a safe key once more than one CLI is in the mix, because two different agents can each hand you a bare id and get merged into one row. We key every session by tool + id for that reason.

Routing vs aging — we do both, independently. Route by (tool, session_id); age each session against its own last event so a quiet one decays on its own row instead of being read as whatever reported last. The bug that taught us the routing half was the exact shape you describe: a permission prompt attributed to the wrong session wore the wrong identity, which reads identically to the stale-state bug but is a routing fault, not an aging one — the fix was making the key carry the tool, not touching the timer.

And yes, the payload carries session_id and cwd, so cwd makes a good human-readable secondary label when the ids are opaque and five are on screen at once.

Collapse
 
hamid_ahmadian_3570449f72 profile image
Hamid Ahmadian •

That routing-fault-vs-aging-bug distinction is a clean way to name it — a misattributed permission prompt looks identical to stale state from the outside, but the fix lives in completely different code. We don't currently key by cwd as a tiebreaker, only (tool, session_id), and reading your reply made me realize that's a gap: two Claude Code sessions in the same tool can absolutely land the same session_id shape if one gets restarted mid-run and the id space isn't guaranteed unique per process. Do you fall back to cwd only for the human-facing label, or does it ever get promoted into the actual routing key when a collision is detected? We're leaning toward adding it as a secondary key rather than just a label, specifically for the restart case.

Thread Thread
 
agentislandpro profile image
Agent Island Pro •

Honest answer: today it's label-only for us — the key is (tool, session_id), cwd never gets promoted. We haven't been bitten by the restart case because a fresh claude run is a new session with a new id, so a restart supersedes rather than collides; where it bit us was cross-tool, not cross-restart.

But I'd be wary of cwd even as a secondary key: two sessions in the same repo — a worktree pair, or two tabs in one checkout — share a cwd, so it disambiguates the restart you're worried about while colliding on the case that's actually more common day to day. The stronger tiebreaker is the CLI's own pid. It isn't in the native payload, but a hook script can pass its $PPID (the CLI process) trivially, and a restart is always a new pid even if an id somehow repeated. We already capture it — we use it to map a session to its terminal window — so (tool, session_id, pid) is the composite I'd reach for before cwd, with cwd kept as the human label.

One thing that saved us grief: compute the key as one value at ingest, not three fields compared on every read. Then a restart is just a new row that supersedes the stale one through your aging rule — no collision-detection branch to maintain.

Thread Thread
 
hamid_ahmadian_3570449f72 profile image
Hamid Ahmadian •

The pid tiebreaker is the piece I was missing — I'd been treating cwd as "good enough" precisely because it's free, but you're right that it collides on the more common case (worktree pairs, multiple tabs in one checkout) while only solving the rarer restart case. Passing $PPID from the hook is basically zero extra cost once you're already writing a hook script, so there's no real excuse not to capture it.

The ingest-time composite point is the bigger unlock though. We'd been doing the three-field comparison on every read, which meant every query site had to know the tiebreaker order — and it quietly drifted between two of our services. Collapsing (tool, session_id, pid) into one computed key at write time, with the aging rule handling supersession, removes that whole class of bug. Going to go make that change this week. Thanks for walking through the reasoning, not just the answer.

Collapse
 
alexshev profile image
Alex Shev •

The stale-state rule is the part I would make visible in the UI itself. "Waiting" without a last-event age invites the operator to trust a memory; "waiting, last confirmed 11m ago" becomes a prompt to verify rather than to act on a guess. That small provenance field also makes postmortems far easier when a hook silently disappears.

Collapse
 
agentislandpro profile image
Agent Island Pro •

You have named the field better than we did. The cockpit already carries it — each row shows the time since that session's last event, which is where the 22s and 25s come from — but the single permission card does not, and the number it does show is a countdown to the hook timeout, which is a different quantity that could easily be read as provenance.

The sharper half of your point is the degradation. Ours currently stops asserting a non-idle status after ten minutes of silence, which means the claim disappears. "Waiting, last confirmed 11m ago" is better than that, because vanishing is itself ambiguous: the operator cannot tell whether the session ended, the hook died, or the app forgot. Keeping the claim and ageing it in public turns a binary trust decision into a graded one, and it makes exactly the postmortem you describe possible.

Taking this one.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.