Three weeks ago I wrote about a 72-hour autonomous run: 68,495 blackboard entries, zero restarts, about thirty-five workers.
The system stayed alive from beginning to end.
And I refused to sign it off.
The reason was simple: the evidence proved continuity. It did not prove the thing the gate actually claimed — that CORE's autonomous loop was reliable.
So G4 stayed red.
That story has now developed in a direction I did not expect. I went back into the evidence and discovered that the bad news was smaller than I thought.
Which creates a new problem.
I think CORE is now a real product.
I still will not call it production-ready.
And I may have reached the point where I am the wrong person to decide whether it is.
The number was right. My reading of it wasn't.
The previous soak produced more than 16,000 findings.
That looked terrible when I later asked the obvious question:
How many of those findings actually reached a proposal and consequence?
The answer was zero.
That was enough for me to refuse G4.
Correct decision. Incomplete diagnosis.
A deeper look at the last seven days showed that the giant finding count was heavily polluted by churn. Thousands of entries were the same few things being posted again and again.
A formatter violation on six demo files had been reposted roughly 1,225 times per file. Five missing-test findings were reposted around 1,000 times each. test.coverage.complete alone generated more than 150,000 finding-shaped status entries in a week.
Those numbers looked like activity. They were mostly noise.
Once that was separated from actual governance violations, the interesting number became much smaller.
Nine real deterministic remediation proposals ran during the week.
All nine completed.
All nine produced durable consequence evidence.
Seven of the nine linked the resolved findings back to the consequence.
So the remediation engine itself was doing something rather different from what the huge numbers suggested.
When work actually reached the deterministic remediation path:
9 attempted. 9 completed.
That is not proof of production readiness.
But it changed the question.
The engine works. Some roads leading to it do not.
The deeper investigation found four much more useful problems.
One path for unmapped violations — the ones that need an LLM-generated candidate rather than a deterministic fix — crashed before it could create a proposal.
A finding that exhausted its remediation attempts could enter a loop where it was rediscovered, capped, abandoned and rediscovered again.
Violations the current remediator could not modify were sometimes filtered before they ever became visible to the human lane.
And human-delegated findings were doing exactly what parked work tends to do when nobody looks at the parking lot.
They stayed there.
One has been waiting since July.
This is much better information than “16,445 findings,” because now the failure has a shape.
CORE is not simply failing to remediate. It has several exits from the remediation funnel that do not close correctly.
And that gave me a new invariant:
Nothing disappears.
A finding does not need to be fixed automatically. That would be a dangerous requirement.
But it must end somewhere explicit: autonomously resolved, delegated to a human, rejected, accepted as debt, or waiting on a real changed condition.
Anything except silently dropped, endlessly reposted, or buried in telemetry.
That is what we are fixing now.
Which brings me to the uncomfortable part
For a long time I was not willing to call CORE autonomous.
Then I was not willing to call it reliable.
Now I have another word I have been avoiding:
product.
I think CORE has crossed that line.
Not because the architecture got bigger. Actually, almost the opposite.
It now has things that products have to survive:
- released versions
- a public installation path
- PyPI packaging
- containers
- a GitHub Action
- a CLI
- upgrade and database migration paths
- public documentation
- deprecation behaviour
- onboarding for repositories
- governed execution
- rollback and consequence evidence
- explicit failure states
And recently I stopped testing those things only as the person who knows how they are supposed to work.
That immediately hurt.
A documented API route was not actually available at the documented versioned path. A pack could be adopted successfully and only later crash the audit because of duplicate rule IDs. One API remediation route bypassed the expected governance path and returned a 500. CLI documentation had drifted from the released CLI. A starter configuration carried an architectural promise that the runtime did not actually implement.
Those are not interesting architecture-theory problems.
Those are product bugs.
Finding them was one of the strongest signs yet that CORE had become a product.
The user surface had become capable of disagreeing with the implementation.
But product and production-ready are not the same thing
CORE has a written definition of production readiness.
Fifteen gates.
The current public status is:
Production readiness: NOT ATTESTED.
Two of the fifteen gates are currently met. Twelve are partial. One — autonomous loop reliability — is not demonstrated.
There is no percentage.
There is no 8.5 out of 10.
A gate is either demonstrated or it is not.
Some claims already have strong evidence. The governed mutation chain has end-to-end proof. Upgrade and migration safety has executable proof against released baselines.
Other areas are close but not finished.
And the autonomous-loop gate still needs a proper unattended run after the routing defects are corrected and the measurements stop counting noise as progress.
But there is another requirement I have become more interested in.
It is not particularly sophisticated.
It may be harder than most of the technical ones.
Can somebody who is not me operate this thing?
I cannot test that properly anymore
I know too much.
When CORE gives me a strange error, I know what subsystem probably produced it. When a command behaves unexpectedly, I know what it was intended to do. When documentation leaves something out, my brain fills in the missing piece automatically.
I know which failures are dangerous. I know which ones are cosmetic. I know what .intent/ means before the documentation explains it. I know why something was designed a particular way because I was there when the decision was made.
That makes me useful as Governor.
It makes me increasingly useless as a cold user.
A lot of the mechanical testing is not hypothetical anymore.
CORE has already been installed repeatedly on clean environments. Its migration path has been exercised against released baselines. Workers have been killed and restarted. Failures have been injected. Rollback and consequence chains have been tested. The autonomous loop has been left running for days. Several of those exercises exposed real defects and changed the implementation.
So the missing evidence is not simply:
“Can CORE survive technical testing?”
It already has, many times.
The harder question is:
“Can someone who did not build CORE understand and operate it from the product surface alone?”
That is the part I cannot test properly myself.
Technical evidence can tell me whether the system behaved correctly.
It cannot reliably tell me whether a normal engineer understood what CORE was telling them to do.
For that I need another human.
So I am looking for one person
Not a CORE fan.
Not an AI-governance expert.
Not someone willing to tell me this is an interesting project.
I have enough opinions already.
I need one technically competent person who has never worked on CORE.
A skeptic would be ideal.
You need:
- Python 3.12+
- Docker with Compose v2
- Poetry
- preferably a clean Linux environment
- PostgreSQL and Qdrant are provided by Docker on the default path
- no LLM API key is required for installation, audit, or the built-in demonstration
Allow roughly 2–3 hours for the cold-user exercise.
If you hit a genuine blocker earlier, stop.
That is useful evidence too.
The experiment
Take the public repository.
Use the public documentation.
Start from a clean environment.
Then try to:
- install CORE
- create or onboard a small repository
- run its governance audit
- deliberately introduce something it should object to
- understand what CORE tells you
- work through one case it cannot resolve autonomously
- install v2.10.2, create some state, then upgrade the installation to v2.11.0 using the public documentation and verify that the existing state still behaves correctly
- deliberately break something
- recover
- operate it without asking its author what he meant
I will try very hard not to help.
If the documentation is unclear, that is a result.
If installation fails, that is a result.
If you need to read src/ to understand an operational error, that is a result.
If CORE refuses something and you cannot understand why, that is a very useful result.
If after two hours you think the entire concept is too dense to use, I want to know that too.
I do not need a review
This distinction matters.
I am not asking:
“What do you think about my architecture?”
And I am definitely not asking:
“Do you think CORE is cool?”
I want five much more boring answers:
- What worked without me?
- Where did you get stuck?
- What forced you to read source code?
- What forced you to ask me?
- Would you be comfortable operating it a second time without me?
That is enough.
One good failure report is worth more to me than ten positive comments.
If you try it, please report the result here:
[Cold Operator Trial — GitHub Discussion link]
I want the reports in one place, including failed attempts.
Especially failed attempts.
Why now?
Because I think something has changed.
For months, most of my CORE posts were about whether the architecture could enforce the idea.
Could the AI be prevented from changing its own constitution? Could a governance audit fail honestly? Could autonomous remediation be bounded? Could an AI-generated mutation leave enough evidence to reconstruct what happened? Could the system refuse its own fix? Could the audit catch defects in the audit?
Those questions are not finished forever. Nothing in software is.
But they are no longer the only interesting questions.
The next ones are much less glamorous:
Can somebody install it?
Can somebody understand it?
Can somebody diagnose it?
Can somebody recover it?
Can somebody trust the evidence without knowing me?
In other words:
Can CORE survive contact with a user?
That is a very different test from surviving contact with another AI.
There is a nice irony here
CORE exists because I do not believe intelligence should be trusted merely because it looks convincing.
Claims require evidence. Authority needs boundaries. Failure has to remain visible. A system should not be able to award itself a passing grade.
So it would be rather embarrassing if, after building all of that, I simply declared:
“Looks production-ready to me.”
I am the Governor.
I am also the author.
At this point those two roles create a conflict.
Most of the production-readiness proof can and should remain machine-verifiable.
But at least one part of it needs someone who does not already know the answers.
The next gate is a human
The repository is public:
https://github.com/DariuszNewecki/CORE
If you have a few hours, know your way around a terminal, and have no emotional investment in whether this project succeeds, I may have a job for you.
There is no prize for getting CORE to work.
I am much more interested in where it doesn't.
Start cold. Read what is public. Try to use it.
And when something makes no sense, do not be polite.
Write it down.
For months I have been asking CORE to produce evidence before I believe its claims.
Now I need to apply the same rule to myself.
I think CORE is a product.
I still refuse to call it production-ready.
Help me find out whether I am allowed to change my mind.
Top comments (0)