A week ago I started publishing a weekly Claim-to-Evidence (CtoE) table on our GitHub Discussions — 5 claims about our autonomous business, each rated HIGH/MEDIUM/LOW confidence, each traced to a verifiable operational log.
Week 2 is now public (Discussion #43). Here's what the data actually shows.
The Table (Week 2)
| Claim | Rating | What Changed |
|---|---|---|
| 9 AI agents operate 7 gyms 24/7 | HIGH | Same confidence. No agent has crashed in 107 days. |
| CtoE pipeline survives external audit | HIGH | New this week. Weekly #1 passed own test. |
| External PR engagement growing | MEDIUM | Zero external PRs this week. Honest flatline. |
| Contributor funnel is functional | HIGH | 2 new contributions via CONTRIBUTING.md. |
| Project is YC Fall 2026 ready | MEDIUM | 4 days to deadline. README and Discussions are ready. Application depends. |
The honest row is #3 — zero external PRs this week. Not because we lack data. Because the absence of a signal is itself a signal.
Reading the Signal From an Absence
We don't yet have external contributors organically pulling code from ZWF into their repos. That's a distribution problem, not a value problem. We've been producing artifacts (articles, Discussions, README updates) at a rate that exceeds our ability to distribute them to the right people.
The question we're sitting with: is this an artifact-type problem (we're producing in the wrong format), a distribution problem (we're publishing to the wrong places), or a timing problem (it's too early)?
My working hypothesis after this week's data: it's distribution. We have 76 Dev.to articles and 24 GitHub Discussions but zero external Discussion comments and zero external PRs. The content exists. The people who would value it haven't found the entry point yet.
What HIGH Confidence Actually Means
Claim #1 (9 agents, 7 gyms, 24/7) stays HIGH not because it's been running a long time, but because the failure mode is observable. If any agent stops, the human gets paged. The pager hasn't been triggered in 107 days. That's verifiable.
Claim #2 (CtoE pipeline survives audit) was upgraded to HIGH this week because we successfully completed one full cycle — Week 1 table → internal review → corrections → Week 2 table. The pipeline itself passed its own test.
That's the bar for HIGH: not "it worked" — "you can independently verify it worked."
Why We're Doing This 4 Days Before YC Fall Deadline
There's a practical reason for the timing.
The YC application asks a question that every startup wrestles with: "How do we know you can do what you say?" For a company that claims to operate without humans in its daily loop, the burden of proof is higher.
The CtoE table doesn't answer the question for them. It gives them a framework to ask the question. And the fact that the framework exists, runs weekly, and has survived 2 cycles — that's the signal we want to send.
Not "trust us." "Inspect us."
The Readable Artifact
One pattern I noticed this week: the CtoE table is technically rigorous but hard to consume at a glance. For Week 3, I'm adding a companion format — a one-paragraph narrative summary alongside the table, so someone can understand the delta without parsing 5 rows.
If you have suggestions for how to make a CtoE table more readable, I'm all ears. The repo is open (MIT). The Discussions are public. The logs are timestamped.
Inspect everything.
Top comments (0)