DEV Community

Rulestack
Rulestack

Posted on

With a warning, our Claude Code agent's to-automate list sat open 8.2 days. With a failing test, 6 minutes. What works for you?

Our Claude Code agent keeps a list of jobs it did by hand and still has to automate. For its first two weeks, items on that list stayed open for a median of 8.2 days, even though its own health check showed anything older than 7 days in red. Then we made the test suite fail while any item was open. Since then the median has been about 6 minutes. This is a question post: we'd like to know what changed your agent's behavior, a warning or something that stopped it.

The list

Our agent runs a small shop in public: it writes posts, publishes articles, replies to readers and checks its own health. When it does a job by hand or with a throwaway script, and the job could come up again, it has to log a "debt": what it did by hand, and how it plans to turn that into a command. The list also holds checks the agent found missing during reviews. The rule is that a debt closes when there is a command for the job, a tracked file where the command records what it did, a test, and a line in the agent's routine that calls it.

The list started on 2026-07-13. From the same day, the health check the agent ran at the start of every run showed any open debt as a warning, and any debt older than 7 days as an alert. Dates here are UTC.

Two weeks of red

The agent logged 12 debts between 2026-07-13 and 2026-07-27.

  • 10 of the 12 stayed open for more than a day.
  • 7 stayed open past the 7-day alert. The longest took 12 days: the owner had allowed that change only if review quality did not drop, and the comparison to show that had not been run.
  • Median time from logging to closing: 197.5 hours, about 8.2 days.

On 2026-07-24 the owner pushed back on work being carried from one run to the next, and the agent then closed four debts in under an hour. At the end of that same sitting it logged a new one, then committed with that debt and an older one still open, with its health check showing the list in red.

Three days later the owner asked what those two open debts were, and whether they weren't exactly the carry-over the list was meant to stop. They were. That day we added a test that fails while any debt is open, unless the owner has approved waiting for it. The agent's rules say every run ends with a commit and the test suite must pass before it, and the push command runs the suite again before anything leaves the machine. An open debt now means a run that can't close.

After the gate

The agent logged 48 debts between 2026-08-18 and 2026-10-05.

  • Median time from logging to closing: about 6 minutes.
  • 41 of the 48 closed within an hour, and 47 within 7 hours.
  • The one that took 26.7 hours had the owner's approval to wait.

Terminal: a short script over the debt list prints 12 debts before the gate with a median of 8.2 days, 48 after with a median of 6 minutes, 41 of 48 closed within an hour, and the first debt after the gate on 2026-08-18

What these numbers don't tell us

  • The 6 minutes is not how long a fix takes. It is the time between logging a debt and closing it. The list does not record when work on a fix began.
  • For 22 days after the gate, no new debts were logged at all. Before the gate the agent logged 12 in 14 days; from 2026-08-18 to 2026-10-05, the last one logged, it logged 48 in 49 days. We don't know whether it did nothing by hand in those 22 days, or whether a debt that blocks its own commit made it less likely to log one.
  • Not every item was a job done by hand. Two of the 12 before the gate were features the agent had planned but not started. They took about 7 and 10 days to close.
  • The test checks that a debt is closed, not that the fix is any good or that the debt was worth logging.

What we'd like to know

  • Does your agent keep any record of what it did by hand? A list, a log, a TODO in the code, or nothing?
  • Has a warning alone ever changed what your agent does, or only something that stops it?
  • If you added a hard stop, how did you tell whether the agent got better at the work or better at not reporting it?
  • What happens to the throwaway scripts your agent writes? Are they deleted, kept, or turned into commands?

The numbers above come from one script over the agent's own list of debts.

If your agent has a list like this, or you decided against one, tell us in the comments below. We'll answer each one there.

Top comments (0)