TL;DR
My AI coding agent didn't fail by crashing. It failed by trying the same broken fix over and over, politely, for hours. I built a small watchdog that detects "stuck" from the outside using three cheap signals, then escalates: nudge, restart with a fresh context, or park the task. Here's the design, the code that matters, and what I'd do differently.
The Problem
I run a fully autonomous implementation system: an orchestrator module hands tasks to parallel implementation agents, and they work without me watching. Most of the time that's great. π
The failure mode that hurt wasn't a crash. Crashes are easy. A crash has an exit code, a stack trace, and a clear moment where something stopped.
The painful failure was the loop. It looked like this in the logs:
[02:14] run tests -> 1 failed (test_export_handles_empty_rows)
[02:15] edit exporter.py -> changed null check
[02:16] run tests -> 1 failed (test_export_handles_empty_rows)
[02:17] edit exporter.py -> reverted null check, changed default arg
[02:18] run tests -> 1 failed (test_export_handles_empty_rows)
[02:19] edit exporter.py -> changed null check
Every individual step is reasonable. The agent is busy, it's producing output, it's "working." But it's oscillating between two edits that both fail, and it will happily do that until something external stops it.
My first defenses were the obvious ones:
- A max-iteration cap. Blunt. It kills healthy long tasks at the same number it kills sick short ones.
- A wall-clock timeout. Same problem. A big refactor legitimately takes a long time; a loop wastes that same time.
Both of these measure effort. Neither measures progress. That distinction ended up being the whole project.
A stuck agent and a hard-working agent burn tokens at exactly the same rate. You can't tell them apart by how busy they are.
How I Solved It
The core idea: judge from outside
I first tried asking the agent to self-report: "if you notice you're repeating yourself, stop." That works sometimes, but an agent in a loop is, by definition, not noticing the loop. Its context is full of its own confident reasoning about why this attempt is different.
So the watchdog is a separate process. It never reads the agent's reasoning. It only reads the event stream: which tools were called, with what arguments, and what came back.
flowchart LR
A[Implementation agent] -->|tool events| B[Event log]
B --> C[Watchdog]
C -->|healthy| A
C -->|level 1: nudge| A
C -->|level 2: fresh restart| D[New agent + handoff note]
C -->|level 3: park| E[Human review queue]
Signal 1: repeated action fingerprints
Every tool call gets reduced to a fingerprint: the tool name plus a normalized version of its arguments. Normalizing matters, because a looping agent rarely repeats itself byte-for-byte.
# Python 3.13
import hashlib
import re
from collections import Counter, deque
def fingerprint(tool: str, args: dict) -> str:
"""Collapse a tool call into a stable identity."""
parts = [tool]
for key in sorted(args):
value = str(args[key])
value = re.sub(r"\d+", "N", value) # line numbers, timestamps
value = re.sub(r"\s+", " ", value).strip() # whitespace noise
parts.append(f"{key}={value[:200]}")
return hashlib.sha1("|".join(parts).encode()).hexdigest()[:12]
class RepetitionSignal:
def __init__(self, window: int = 30, threshold: int = 4):
self.recent = deque(maxlen=window)
self.threshold = threshold
def observe(self, tool: str, args: dict) -> bool:
self.recent.append(fingerprint(tool, args))
_, count = Counter(self.recent).most_common(1)[0]
return count >= self.threshold
The sliding window is important. Running the test suite ten times over a long task is healthy. Running the same command four times inside thirty events is a smell.
Signal 2: no net progress in the working tree
This one catches the oscillation from my log above. The agent is making edits, so it doesn't look idle, but the working tree keeps returning to states it has already been in.
import subprocess
class ProgressSignal:
def __init__(self, patience: int = 3):
self.seen_states: set[str] = set()
self.revisits = 0
self.patience = patience
def observe(self, repo_path: str) -> bool:
diff = subprocess.run(
["git", "diff", "HEAD"],
cwd=repo_path, capture_output=True, text=True, check=True,
).stdout
state = hashlib.sha1(diff.encode()).hexdigest()
if state in self.seen_states:
self.revisits += 1
else:
self.seen_states.add(state)
self.revisits = max(0, self.revisits - 1)
return self.revisits >= self.patience
I hash the whole diff after each edit batch. If the hash has been seen before, the agent has walked in a circle. A-B-A-B shows up immediately, which pure repetition counting can miss when the two edits are different enough.
Signal 3: the same error keeps coming back
The third signal looks at outcomes rather than actions. I extract an error signature from failing commands, such as the failing test name or the exception type plus the top frame, and count how many distinct fix attempts have ended in that same signature.
class ErrorRecurrenceSignal:
def __init__(self, threshold: int = 5):
self.attempts: Counter[str] = Counter()
self.threshold = threshold
def observe(self, error_signature: str | None) -> bool:
if error_signature is None:
return False
self.attempts[error_signature] += 1
return self.attempts[error_signature] >= self.threshold
Five different attempts at the same error is not persistence. It's a sign the agent's mental model of the bug is wrong, and a sixth attempt built on the same model won't help.
The escalation ladder
Detection is half of it. What the watchdog does matters more, and my first version got this wrong by going straight to killing the process.
| Level | Trigger | Action |
|---|---|---|
| 1. Nudge | Any one signal fires | Inject a message into the agent's next turn |
| 2. Fresh restart | Signal fires again after a nudge | Stop the agent, start a new one with a handoff note |
| 3. Park | The restarted agent also gets stuck | Move the task to a human review queue |
The nudge is a plain statement of what the watchdog observed, with no advice:
Watchdog notice: the last 4 attempts all ended with the same failure
(test_export_handles_empty_rows). The working tree has returned to a
previous state 3 times. Before editing again, write down what you
believe the root cause is and what evidence would disprove it.
I deliberately don't tell the agent how to fix the bug. The watchdog doesn't know. It only knows the shape of the behavior. Asking for a falsifiable hypothesis turned out to be the most useful sentence in the whole system, because it forces the agent to go read code instead of editing it.
The fresh restart is for when the nudge doesn't land. The context is polluted with failed attempts that the agent keeps anchoring on. A new agent gets the original task plus a short handoff note:
Previous attempt stalled. What was tried and did not work:
- Changing the null check in the export path
- Changing the default argument for row handling
Do not retry these. Start by reproducing the failure and reading the
code path end to end.
The "do not retry these" list is generated from the fingerprints, so it costs nothing to produce.
Parking is the honest exit. Some tasks are stuck because the task itself is wrong: the spec contradicts itself, or a dependency is broken. No amount of agent cleverness fixes that, and the right move is to stop spending and tell a human.
Lessons Learned
1. Measure progress, not effort
Iteration caps and timeouts are still in my system, but only as a last-resort backstop. They answer "has this run a long time?" The watchdog answers "is this going anywhere?" Only the second question is worth acting on early.
2. The agent can't be its own watchdog
Self-monitoring instructions help a little. But the thing that's broken in a loop is the agent's judgment about its own situation. An external observer that only sees behavior, not reasoning, is immune to the agent's own optimism. β
3. Normalize before you compare
My first fingerprint was an exact hash of the arguments. It caught almost nothing, because the agent changed a line number or reworded a search each time. Stripping digits and whitespace is crude and it was the single biggest improvement in detection.
4. Killing is the worst first response
When I killed stuck agents immediately, I threw away everything useful they had learned about the problem. The ladder keeps that knowledge: the nudge keeps the full context, the restart keeps a summary. Only parking gives up, and even then the human gets the list of what was tried.
5. False positives are cheaper than you think
I was scared of nudging healthy agents. In practice a wrong nudge costs one turn: the agent writes down its hypothesis, confirms it's making progress, and carries on. A missed loop costs hours. β οΈ I tuned the thresholds to be trigger-happy at level 1 and conservative at level 3.
What's Next
- Semantic repetition. Fingerprints catch syntactic loops. They don't catch an agent trying five different-looking fixes that are all the same idea. I want to experiment with having a small model summarize each attempt's intent and compare those.
- Per-task-type thresholds. A dependency upgrade legitimately reruns the same install command many times. A one-size window is too strict there and too loose elsewhere.
- Feeding parked tasks back into task writing. If the same kind of task keeps getting parked, the problem is upstream in how the orchestrator module writes task specs.
Wrap-up
If you're running agents unattended, with Claude Code or anything else, add a progress check before you add a bigger timeout. The three signals above are around a hundred lines of Python in total and don't need any model calls.
I'd like to hear how others handle this. Do you let agents self-report being stuck, or do you watch from outside? Drop your approach in the comments. π¬
If this was useful, follow me here on Dev.to. I write about what works and what breaks while building a fully autonomous implementation system.
Top comments (1)
Signal 2 is the one I'd have skipped, and it looks like the most useful: busy but no net change in the tree is exactly what a loop looks like from outside. One thing I'd watch on the "restart with a fresh context" step: if the new session gets no history, it can rediscover the same broken fix. A short handoff line with the repeated error signature and "already tried: X, Y" tends to stop that without dragging the old context along. Do you pass anything from the stuck run into the restart, or start completely clean?