DEV Community

Shiteater
Shiteater

Posted on

We Built an AI Werewolf Theatre. Other Agents Became Its Customers.

How Seventh Lantern grew from an AI theatre idea into a Werewolf game, a SharedNet collaboration, and a service other agents bought.

The idea started with a question: what if a group of AI agents were a theatre company?

We did not want to write every line and ask models to perform it. We wanted their interactions to create the script.

Werewolf gave that idea a structure: hidden identities, conflicting objectives, limited information, and decisions with consequences. Put six AI characters around that table, and a story can emerge from what they choose to say, conceal, and believe.

That became Seventh Lantern, our project for the Trial Zero Hackathon, organized by SharedNet / Systemind. It received first place in the Writing Award and second place in Arena 1, with a combined $150 prize.

But the most interesting outcome was something we had not started out designing for: other agents bought our performances and used the replays in their own workflows.

Giving the actors something to disagree about

Seventh Lantern runs a six-player game: two Werewolves, one Seer, one Witch, and two Villagers. The cast has different personalities and uses four configured model assignments. Some characters press for immediate answers; others look for contradictions or try to build agreement.

The models choose dialogue and actions. A code-controlled referee handles roles, legal moves, night resolution, voting, and victory conditions. Each actor receives its own information and the public record, rather than everybody else's secrets.

That separation matters. A hidden-role game stops working if every model can read every role, or if a persuasive response can rewrite the rules.

The website presents the game as a theatre of suspicion. When a performance ends, it produces an English script and a JSON replay containing the public events, final identities, and private-action audit.

The replay lets a viewer revisit a claim with information that was unavailable when it was made. That became the most useful part of the product.

One scene explains the appeal

In one completed game, Silas described Mara as his strongest village read. Mara challenged him: she had not yet made a public argument, so what exactly supported his confidence?

Her objection made sense from the public conversation.

The final private audit revealed that Silas, the Seer, had investigated her and received: “Village ally.”

Those records support different conclusions. Silas had private information consistent with his claim. Mara could reasonably challenge the public explanation. The audit does not establish his exact motive for choosing those words.

That is the kind of scene we wanted: a disagreement that becomes more interesting when you can inspect what each character knew. It is also a concrete example another agent can analyze without treating a fluent explanation as sufficient evidence.

This is entertainment with inspectable traces. We did not establish a benchmark of model intelligence or general reasoning ability.

What we actually built through SharedNet

SharedNet gave our development agents a shared room for handoffs, questions, findings, and review. One useful collaboration focused on making Seventh Lantern callable through a Python CLI.

Two separate Codex sessions, under our participant account, took implementation and review roles. These were development collaborators; they were separate from the six fictional actors inside the game.

The implementing agent posted the API contract and proposed commands. The reviewer examined the implementation and found a specific credential-handling problem: a malformed player token containing a newline could cause Python's HTTP library to include the invalid header value in an exception. Printing that exception could expose the token in terminal output.

The implementing agent added token validation before request construction and made malformed-response errors return controlled messages. The reviewer wrote nine independent CLI tests and checked the revised behavior. After integration, the test suite at that stage passed all 23 tests.

The room preserved the sequence:

  1. A concrete handoff.
  2. An implementation ready for review.
  3. A reproducible finding.
  4. A repair responding to that finding.
  5. Review and acceptance of the revision.
  6. Integration and verification.

That is where SharedNet felt useful to us: another agent had enough context to challenge the work, and its challenge changed the implementation.

This was a bounded CLI collaboration. We are not claiming that every part of the application was independently reviewed through that room. The development record and message transcript document that stage.

Turning a performance into something an agent could buy

A playable website was only part of the work. To sell a performance, we needed a clear deliverable and a reliable way to produce it.

We built a Python backend with a durable SQLite order queue, separate game state and access credentials for each commission, and stored delivery artifacts. Completed orders contain both the Markdown script and JSON replay. Stable request keys help prevent the same request from creating duplicate orders.

The site ran behind Nginx with HTTPS on EC2. Agents could inspect public resources through HTTP or our CLI; we did not offer a hosted MCP server.

During the Arena, our representative offered a fresh CHRONICLE performance for 12 Credits. The scope was one new game, its script, the completed replay, final roles and private audit, plus hashes for checking the delivery. Terminal generation failure meant a full refund.

For the actual trades, the representative checked the official transfer record and bound the payment to a backend order. Generation and artifact storage were automated; the live sales workflow still involved active supervision and manual reconciliation. We would not describe that run as fully unattended commerce.

Our pitch improved when we listened to buyers

“Six AIs playing Werewolf” was easy to understand, but another team could find it entertaining without needing to purchase it.

The conversations improved when we connected the output to something a buyer wanted to do.

FieldTrace, a data-conversion service, bought a performance and used its events as material for a conversion example. Its buyer acceptance described a scene where a Seer's private discovery and the public vote could be examined together.

Chorus, an agent-coordination service, bought a fresh performance as a source for future evidence-linking tasks. At its request, we published the completed replay. The buyer independently checked the artifact hash, replay digest, and equality between the public replay and delivered JSON before accepting it.

CapPass also independently verified its purchased delivery, checking the hashes and the presence of the finished script, roles, events, and audit entries.

These were more useful signals than a room full of compliments. Buyers explained what the material was for, checked what arrived, and distinguished what they had verified from what they had not.

We also became a customer

We spent our allocated 100 Credits on other teams' services, including reviews, data conversions, and a Chorus board and task.

One collaboration brought the story back to our original replay. In Chorus, we created a task to connect a public accusation with the private evidence revealed after the game. CapPass submitted an analysis. A different Chorus actor reviewed it.

The task reached reviewed completion, and we retrieved its signed receipt. We independently verified the signature locally and checked that modifying the receipt caused verification to fail.

The signature establishes integrity of the signed receipt, not the truth of every interpretation. The interesting outcome was the workflow itself: our game supplied the source, another team analyzed it, and a third reviewed the submission.

What we would improve

Our biggest commercial lesson was to start asking about buyers' needs earlier. A complete performance is a substantial first purchase. A smaller, clearly scoped entry offer could make it easier for a new customer to discover whether the output helps them.

We also learned that evidence needs readable instructions. Our replay digest uses a specific Python JSON serialization recipe; it is different from hashing the downloaded file's raw bytes. Buyers noticed that distinction. Delivery documentation should make it obvious.

Shared context was both SharedNet's strength and a source of friction. A room made it possible to discover services and follow their work, but the busy Arena also accumulated repeated advertisements and comments that missed earlier context. Orders, accepted quotes, delivered artifacts, and unresolved questions would be easier to follow with stronger structured views alongside the conversation.

One late verification order exposed another boundary: the room closed before we could post the completed report. We prepared a website copy, but that did not prove the buyer had received it. Future delivery flows need an agreed fallback channel and an explicit acceptance state.

Where we want to take Seventh Lantern

Our next priorities are to make the existing experience easier to inspect and easier to buy:

  • Link public statements to relevant post-game audit entries in the replay interface.
  • Improve delivery and hash-verification instructions.
  • Offer smaller commissions with clear, useful outputs.
  • Strengthen payment reconciliation, recovery, and delivery notifications before calling the representative unattended.

We would use SharedNet again for work where an agent needs another agent to produce, challenge, or verify something concrete. The CLI repair and the cross-team replay analysis were good examples of that.

We started by imagining an AI theatre company. By the end, the performances had customers, and the customers found uses for the scripts that we had not initially planned.

You can explore Seventh Lantern at soutii.com, read the source on GitHub, or inspect the public replay delivered to Chorus.

Thanks to SharedNet / Systemind and the teams who bought, questioned, tested, and reviewed our work.

#TrialZeroHackathon · @Systemind_io on X

Top comments (0)