We just closed a preregistered cohort of 30 real development tasks executed by a paired-agent workflow: one agent implements, a second agent from a different harness independently recomputes the tests and records process events before anything ships. The cohort was small, descriptive, and single-operator by design. No causal claims. The point was to learn what actually happens when you write governance rules down and then do real work.
The headline numbers look excellent:
Recorded slots 30/30
Complete required chains 30/30
Gate coverage 30/30
Verification events 111
Export gate ready
Three of the most interesting numbers are the ones that look good but should not.
- Reservation timing collapsed under real work The preregistration said: reserve a slot before work begins, never backfill completed tasks. Real life disagreed three times. In three later slots, implementation commits landed before the slot was reserved. On one dense day, 17 commits of real work shipped with no slot at all — completed tasks that, per our own rule, can never be backfilled. The mechanism is boring: pressure arrives, work starts, and "open a slot first" is an extra step that nobody remembers at 2 a.m. Writing the rule down did not change the behavior. The honest recording is that reservation discipline degraded exactly when task density went up.
- Rework = 0 is a blind spot, not a virtue The final report says Rework incidence: 0/30. That is false in spirit. During the cohort we redid a slot closure twice after an infrastructure incident wiped local commits, and plenty of fixes went red-test-first before going green. The schema records rework only when someone voluntarily writes a rework event. Nobody volunteered. This is the absent-zero error wearing a governance costume: an absent record rendered as a confident zero. We now write that rule down explicitly — an absent event must be reported as unknown, never as zero — in docs/validity-states.md.
- Duration coverage drifted mid-cohort Only 9 of 30 slots carry explicit durations. The early slots predate the duration rule; the later ones record it. So duration coverage improves over time because the recorder changed, not because the process got better. Any "trend" in that column is a measurement artifact. What v2 must do differently
- Machine-check reservation timing. If the first commit predates the slot reservation, that is an automatic violation event — not a paragraph in a handover note.
- Make rework mandatory. A failing verification that is followed by a correction must emit a rework event. Voluntary rework reporting produced zero.
- Freeze the recorder before slot 1, including duration and rework fields.
- Add at least one more operator. A single-operator cohort cannot separate discipline from personality. The gate said "ready". We chose not to export. The export gate returned ready: manifest valid, safety scan clean, source registry verified. We are still not exporting the event data. Thirty production-adjacent tasks from one operator is a small, non-random sample where cross-field reidentification risk is highest and learning value for outsiders is lowest. That is the gate working correctly. A gate that can only say "no" is theater. A gate that can say "ready" — while a human still says "not worth the risk" — is an actual decision. The durable output is the negative result and the rules it produced, not the dataset. As before: descriptive only, no causal claims, and the full trust layer remains reproducible at janzong/agent-lab-trust. If you run agents on real work, the uncomfortable question is not "did my rules pass?" It is "which of my zeros are just absences?"
Top comments (0)