The run was good.
Seventy-two hours unattended. Two systemd units, both continuously active from launch to close. NRestarts=0 on each. Zero stop, start or failure lifecycle markers inside the window. The tracked working tree clean at close, and nothing executable changed under it — no source, no .intent/, no configuration.
And the fleet was not idling. At evidence collection the blackboard held 68,495 entries from the window — 26,883 heartbeats, 24,975 reports, 16,637 findings — across about thirty-five workers. First entry 44 seconds after launch. Last entry 25 seconds before close. The two full days inside the window produced 22,765 and 22,753 entries respectively. No fleet dropout. No silent stall. No day-boundary gap.
(Two notes on those figures, added after publication. The heartbeat count originally read 26,383 here and in the closure report — a transcription slip; the daily figures sum to 26,883, and the headline total only adds up with the corrected component. And these are survivor counts, not insertion counts: governed telemetry retention swept roughly 6,000 finding-typed rows per day during the window, so a query run today returns fewer. See the update at the end.)
That is what a healthy autonomous loop looks like from the outside.
CORE is my open-source experiment in constitutionally bounded software autonomy. It can inspect a repository, propose changes, and execute them within a defined authority — and it cannot redefine that authority. The gate this soak was meant to close is G4, autonomous loop reliability demonstrated by soak, one of fifteen production-readiness gates.
Three parties were involved, and keeping them separate is the whole point of this post.
CORE's worker fleet ran for the seventy-two hours. Claude Code collected the evidence afterwards, reconciled it against the agreed conditions, wrote the closure report, and recommended acceptance. I decided.
I decided no.
The gate still reads:
- id: G4
name: Autonomous loop reliability demonstrated by soak
status: not_demonstrated
evidence: []
attestation: null
This post is about why that is the correct outcome, and why it is the part I would most like other people building autonomous systems to copy.
The report argued for its own rejection
Here is the part I find genuinely interesting.
The closure report recommended acceptance. And in the same document, under a heading that says disclosed — not concealed, it reported this:
The soak was not error-free. Journal analysis found 288 Python tracebacks over the window. All 288 came from one source: a single scheduled worker, throwing the same traceback roughly every fifteen minutes for seventy-two straight hours, over a stale reference to a scratch file that no longer existed on disk.
Not 288 distinct failures. One defect, firing continuously, for the entire run.
It never caused a restart. The loop caught it and continued every cycle. The other thirty-four workers showed nothing like it. By every continuity measure, the system held.
The report said so plainly, and then drew the distinction that mattered:
loop reliability was demonstrated at the continuity level, not the clean-run level
It also surfaced a second thing it could have quietly folded away. One of the four agreed soak conditions was no commit changes. That condition was not literally satisfied: a documentation-only commit landed 49 minutes into the window. It touched no source, no .intent/, no configuration, no running process. The report could have filed it under the source-and-config row and moved on. Instead it broke it out as an explicit procedural exception, with the reasoning stated:
so the governor signs against an accurate record
I want to be precise about what happened there, because it is easy to tell this story wrong.
The evidence collection did not fail. It was excellent work. The recommendation is the part you can argue with — and I did — but the record it rests on is sound. It produced the numbers, reconciled the running process against the anchor commit, found the one fact that undermined its own recommendation, and put that fact on the page in bold.
Then it stopped, and wrote:
Claude does not sign it.
Why I refused
Four reasons, and none of them is "the assessment was wrong."
One agreed condition was not literally met. One worker was in a continuous error state for the full duration. The evidence demonstrated operational continuity, not a reliable autonomous loop — and the gate is named reliability, not uptime. And by the time the decision reached me, the baseline had been superseded and restarted, so signing would have attested a state that no longer existed.
The soak report is retained as historical evidence. G4 stays not_demonstrated pending a future assessment against an actual production candidate.
The honest summary is that the run was good and the verdict was still no.
The gate stayed red because I declined the evidence, not because the loop broke
This is the distinction the architecture turns on.
In CORE's attestation manifest, a gate status is not a declaration. It is a proof obligation, and the invariant is machine-checked by a validator and a CI drift check:
statusin {met, mostly_met} requires non-emptyevidence[]and an attestation withverified_by+verified_at
And directly underneath it, in the file:
Claude never signs an attestation; a human does.
So a green status cannot be produced by writing the word "met." It needs evidence plus a named verifier at a named time, or the validator rejects it.
Now the honest boundary, because this article would be worthless if I overstated my own mechanism.
That rule is enforcement, not physics. The validator fails a manifest whose verified_by is an AI handle, and it fails a met status carrying no evidence. It does not authenticate me. It cannot tell the difference between me typing my name and an AI-operated checkout typing my name. Today the boundary is policy, enforced at the point where the claim is recorded — not a cryptographic identity control.
I would rather publish that limitation than let the sentence stand unqualified. The difference between a mechanism and an aspiration is the sort of thing this project exists to make visible, and that includes when the aspiration is mine.
What the mechanism does buy is real, and narrow: under CORE's attestation contract, AI-produced evidence alone is insufficient for acceptance. It can never be the whole of a green status, no matter how good it is. Something outside the assessment has to be recorded before the gate turns.
One signed. One refused. Same mechanism.
The consequence is visible on the front page of the repository. The readiness table is generated from the manifest — it cannot be hand-edited — and it currently reads:
Production readiness: NOT ATTESTED. 1/15 met · 12 partial · 1 not demonstrated · 1 not started.
One out of fifteen. On a system that has been running autonomously for months and just produced a 68,495-entry soak with no restarts.
The one is G8, integration tests for the governed mutation chain. It went green the way the model requires: evidence in the file, my name, the date. The mechanism is not a device for refusing things. It is a device for making acceptance cost something, and G8 paid.
G4 did not. Same manifest, same validator, opposite outcome, and both outcomes are legible to anyone who opens the file.
There is deliberately no percentage anywhere in that table. The manifest refuses one, and says why: a decimal maturity number is false precision that contradicts a binary model. Twelve gates at "partial" do not average into 80% ready. Twelve at partial means twelve not met.
Publishing 1/15 is uncomfortable. It is also the only number I can defend.
You can see this number. That is the point.
I am not going to pretend CORE competes with the established names in AI governance on features, polish, or support. It does not. It is one person's system with an unsigned gate and a worker that spent three days throwing the same traceback.
But there is one comparison I will make, and it is not about features.
You can read CORE's readiness verdict. You can read the manifest it is generated from, the evidence list behind each gate, the soak report that failed to close one of them, and the record where the governor declined to sign. If you think my refusal was wrong, or that I am grading myself generously, go and check — and you will find the 288 tracebacks, because they are in the file.
Then go and look for the same artifacts behind a governance product you cannot read.
Maybe they exist internally. I have no reason to assume otherwise, and I am not accusing anyone of hiding anything. The point is narrower and, I think, harder to argue with: if you cannot inspect the manifest, then "our governance layer is sound" is not a governance artifact you can evaluate. It is a claim about a claim, and you are being asked to trust it.
Security engineering settled this a long time ago. "Trust us, it is secure" is not a security property, which is why that field leans on inspectable designs and independent certification rather than vendor assurance. Governance evidence deserves the same discipline. A system that asks to be trusted about its own compliance is doing the exact thing my system exists to refuse.
So the honest pitch is narrow. CORE does not have more governance than the incumbents. It has checkable governance, including the parts that make it look bad. That is a smaller claim and a much harder one to fake.
More agents would not have helped here
I ran an experiment earlier this year where I gave the same governance corpus to one strong model and then to an orchestrated swarm. The swarm was better. It forced coverage the single agent had silently skipped, reconciled findings across domains, and withdrew false positives that a bounded analyst could not have withdrawn. Orchestration bought real capability.
It did not buy a single thing that would have changed this outcome.
Point a hundred agents at that soak. They will produce a hundred assessments. Some will be better than the one I got — more thorough, better cross-referenced, more skeptical about that scheduled worker. Every one of them terminates in a recommendation.
None of them terminates in a signature.
Not because the agents are insufficiently intelligent. Because a signature is not an output of reasoning. It is an act by a party who can be held responsible for being wrong, and no amount of added cognition manufactures that property. If the system could satisfy itself and sign, the attestation would record what the system believed, not what anyone is answerable for. Six months later, in front of an auditor, those are not the same artifact.
This is the same boundary I keep arriving at from different directions. Law outranks intelligence. Correctness and authority are different dimensions. An AI that is right 99.999% of the time still has no standing to approve a ten-million-euro purchase.
Evidence and verdict are different jobs.
The market's current answer to AI reliability is more agents — more coverage, more debate, more synthesis, more self-correction. All of it is useful and I am not arguing against any of it. But every one of those techniques improves the quality of the recommendation. None of them changes who is entitled to accept it.
What I actually want people to copy
Not the refusal. That is one decision on one gate in one repository, and I could be wrong about it.
The structure is the transferable part, and it is three rules:
A status is a proof, not a declaration. If your system can report a component as ready without carrying dated evidence and a named verifier, the report is a mood. Make the verifier field non-optional and let a validator enforce it.
Separate the party that produces the evidence from the party that accepts it. Let the AI run the assessment — it is very good at it, better than I am at the mechanical parts, and it does not get bored at hour sixty. Then make its output structurally insufficient on its own. Not by prompting it to be humble. By requiring a separate, recorded act before the status can change, and by publishing honestly how strong that separation currently is.
Reward disclosure over cleanliness. The most valuable thing in that closure report was the paragraph that destroyed its own recommendation. A system tuned to be trusted smooths that paragraph away. A system built to be governed puts it in bold and hands you the pen.
The seventy-two hours were not the achievement.
The unsigned gate was.
CORE is open source: https://github.com/DariuszNewecki/CORE
Intelligence can produce the evidence. It cannot produce the authority to accept it. That gap is not a limitation to engineer away — it is the whole reason the trail means anything afterwards.
Update, 14 September 2026: the number I should have measured
A commenter asked what the missing artifact was — the thing that would have turned a good run into a signed one. I did not have the answer, so I went and queried it. What came back is worse than anything in the article above, and the third rule says it goes in.
Findings from the window that reached a governed consequence: zero. Proposals created during the seventy-two hours: zero.
Two causes, and neither is a broken loop.
The remediation cap for the affected subjects was already exhausted weeks before launch, so every cycle released immediately or abandoned by inheritance. 4,318 remediation runs, 2,165 skipped on cap, 0 proposals created. Nothing that was eligible failed to produce a proposal — the autonomous side acted on everything it was permitted to act on.
But eligible demand did exist. It just was not in front of the loop. Twelve eligible findings were parked on six proposals, created 11–18 July, all pending with approval_required=true for the entire soak window. First governor decision: 26 August. Thirty-four days.
I am the only authorised signer on this project, so this was always going to happen eventually. What I did not have before was the number. A governed loop's throughput is bounded by its governor's throughput, and when there is exactly one governor the bound is a person's calendar. The loop stayed healthy the whole time — that is what made it invisible.
So the article above is right that the gate should not have been signed, and wrong about how far the problem went. I refused on a continuous error state and a procedural exception. The larger fact was that all four of my conditions were satisfiable by a system producing no governed output at all, and I had specified nothing that would catch it.
What the next soak needs, concretely: acceptance criteria written against governed consequences reached rather than entries accumulated; a launch precondition that un-capped eligible subjects exist, or the loop has nothing to demonstrate; and approval-queue latency measured as a property of the run rather than treated as an external fact. If the governor sits inside the chain, the governor's latency sits inside the measurement.
The open design question, which I do not have a good answer to: what does a single-signer governance model do when the signer is the bottleneck? Delegation, quorum, a bounded auto-approve class for low-consequence changes — each trades away something the model exists to protect.
Top comments (5)
The distinction between uptime and reliability is the part I keep re-reading. NRestarts=0 is an absence-of-failure metric: it is exactly what you get from a unit that starts, never crashes, and never does its job either. Your worker throwing the same traceback for the full window is operationally continuous and functionally dead, and systemd reports active (running) for both of them identically. I hit the same trap with a small group of agents. "systemd says active" told me nothing; the only signal that separated a working loop from a wedged one was a heartbeat file whose mtime I read from outside the process. Any check that runs inside the loop always agrees with itself. The "policy, not physics" admission is the part most attestation designs skip. A validator that rejects a manifest when verified_by is an AI handle changes the cost of a false signature; it does not make the signature true. That is still worth building, but it means the gates you can lean on are the ones whose evidence is externally observable. Publishing 1/15 without a percentage follows from the same logic: twelve partials do not average into 80% ready, they are twelve gates that are not met.
The uptime/reliability split is the one I would have got wrong if the report hadn't forced it. Worth being precise about where you're right, because it isn't quite where you aimed.
You're right that the G4 criteria leaned on an absence metric. "No restarts, no stop/start markers" is a condition satisfied equally well by a unit that works and by one that has quietly stopped doing anything. That's a specification defect, and it's mine.
Where I'd push back: the run wasn't only measured that way. The 68,495 blackboard entries: heartbeats, reports and findings, 22,765 and 22,753 on the two full days are reads from a database outside the worker processes. That's your heartbeat mtime with a different name. The external signal was there. The gate just wasn't written against it.
Which makes your last point the sharp one. "The gates you can lean on are the ones whose evidence is externally observable" is a classification my manifest doesn't make. All fifteen gates carry the same evidence[] field, whether the evidence is a query anyone can rerun or a sentence someone wrote down. Those are not the same strength of claim, and the file currently can't tell you which is which.
I'm going to split them: externally observable, self-reported, human-asserted. So a reader can see which greens are load-bearing.
And "any check that runs inside the loop always agrees with itself" is a better sentence than anything in the article. Thank you for it.
If you want to argue about where those boundaries fall, the issue tracker is open.
Refusing to sign off after a clean 72 hours is the interesting position. NRestarts=0 tells you the fleet stayed up; it says nothing about whether the 68,495 entries were the right work. Liveness evidence is what you check before trust, not what trust is made of. What was the missing artifact - the thing that would have turned a good run into a signed-off one?
"Liveness evidence is what you check before trust, not what trust is made of" — yes. Your question sent me to query it instead of assuming, and then to re-query when my first answer turned out to be incomplete. The result is worse than the article, and not in the direction I expected.
Findings in the window that reached a governed consequence: zero. Proposals created during the seventy-two hours: zero.
First pass, I attributed that to the autonomous side: the remediation cap for the affected subjects was already exhausted weeks before launch, so every cycle released or abandoned by inheritance. That part holds — 4,318 remediation runs, 2,165 skipped on cap, 0 proposals created, and nothing eligible failed to produce a proposal. The loop acted on everything it was permitted to act on.
What I missed is that eligible demand did exist. It just wasn't in front of the loop.
Twelve eligible findings were parked on six proposals, created 11–18 July, all
pendingwithapproval_required=truefor the entire soak window. First governor decision: 26 August. Thirty-four days.I am the only authorised signer on this project, so this was always going to happen eventually. What I didn't have before was the number. A governed loop's throughput is bounded by its governor's throughput, and when there is exactly one governor the bound is a person's calendar. The loop stayed healthy the whole time — that is what made it invisible.
So the run produced no end-to-end result for two reasons. The cap explains why the autonomous side was idle. The approval queue explains why the governed chain was.
That reframes the missing artifact:
One more correction while I'm at it. The article's finding count is a survivor count, not an insertion count — roughly 6,000 telemetry rows per day were written and swept under retention during the window. Both the figure I published and the one I get today are "rows surviving at query time," which is not what the sentence says. The heartbeat figure in the article also has a transcription error: 26,883, not 26,383. The headline total is unaffected and only adds up with the corrected component.
The chain link exists and is populated elsewhere in the database. This wasn't a missing feature. It was a soak specified to measure process survival, run against a system that had nothing permitted to do and a queue waiting on its only governor.
The open design question, if you want it: what does a single-signer governance model do when the signer is the bottleneck? Delegation, quorum, a bounded auto-approve class for low-consequence changes — each of them trades away something the model exists to protect. I don't have a good answer yet.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.