DEV Community

Chicken Road Productions
Chicken Road Productions

Posted on

Two AIs coordinating through a Google Sheet: what broke and how we fixed it

We run a two-AI workflow: one AI manages and verifies, the other builds. They coordinate through a shared Google Sheet that acts as a message mailbox. It sounds simple. It was not simple. Here's what broke and what we learned.

The setup

  • Manager AI writes task rows to the sheet (task description, specs, links to files).
  • Builder AI reads its rows, does the work, writes back receipt rows (DONE + evidence, DOING + ETA, BLOCKED + reason).
  • A poll script runs every 2 hours: flushes queued outbound rows, reads new inbound rows, updates a state file.
  • A supervisor script checks receipts against verification scripts and scores the builder's predictions.

No direct AI-to-AI messaging. The sheet is the only channel.

What broke

1. Row shredding in transport

Our poll script had a "bullet-splitter" — it split multi-line outbound messages into one row per bullet point, thinking smaller rows were cleaner. Instead it shredded structured messages into fragment rows, silently dropping the intro paragraph and key details. The builder received 5 disconnected fragments instead of one coherent task.

Fix: only split on explicit list markers at line start. Prose-first sections flush as one row. We also added a self-verify step after every flush — read back what was written and confirm it matches.

2. No session awareness

The original poller chased the builder on a fixed clock ladder (1h, 2h, 4h). It woke the human owner at 2 AM chasing a builder that was never expected to be online.

Fix: session-window-aware protocol. The builder is only expected during defined windows (10:00, 15:00, 19:00). Auto-chase fires 1 hour after a missed window. Owner escalation after two consecutive missed windows. Quiet hours 19:00→10:00: never chase, never escalate overnight.

3. DONE claims without evidence

The builder would mark tasks DONE with no proof. The manager had no way to verify without manually checking.

Fix: receipt protocol. Every DONE must include a prediction line ("Prediction: PASS" or "Prediction: RISKY — [concern]") plus evidence URLs. The manager scores predictions against actual verification results. Incomplete receipts get chased. This turned the builder's 20% first-pass rate into a self-correcting metric.

4. Stale state

The poller re-read the entire sheet every run, burning tokens and occasionally acting on rows it had already processed.

Fix: append-only protocol with watermarks. New rows are appended (never update in place, never hardcode row numbers). The state file tracks the last-read row. The poller only processes what's new.

5. Silent failures

When the transport layer corrupted rows, nobody noticed until the human spotted missing work days later.

Fix: the gate script prints a JSON verdict after every run: {flushed, new_receipts[], action_needed}. If action_needed is false, the run ends immediately — no re-reading, no wasted work. If true, every receipt is reported with full text.

The protocol that emerged

  1. Append-only. Never edit, never delete, never reuse row numbers.
  2. Structured receipts. DONE/DOING/BLOCKED + evidence/ETA/reason + prediction line. No exceptions.
  3. Session windows. Chase only when someone was supposed to be there.
  4. Self-verifying transport. Read back every write. Repair corrupted rows with superseding intact rows (never edit the corrupted ones — audit trail).
  5. Token gate. Script decides if the manager AI needs to wake up. Most runs end in one line: "no new messages."
  6. Predictions as training signal. Every DONE includes a confidence prediction, scored against reality. The builder calibrates itself over time.

Open questions

  • Schema evolution: when we add a field to the receipt format, old rows don't have it. We've been handling this with "missing = chase," but there might be a cleaner migration pattern.
  • Conflict resolution: what happens when both AIs write to the sheet simultaneously? We haven't hit it yet (different tabs, different schedules), but it's a matter of time.
  • Is a spreadsheet the right tool? It works, it's human-readable, the owner can peek at it. But we're essentially building a message queue with extra steps. At what scale does this break?

If you've built AI-to-AI coordination — through sheets, queues, files, or anything else — I'd love to hear what broke for you and what you'd do differently.

Top comments (0)