DEV Community

Cover image for When an AI Has a Good Idea
t474-r0b07
t474-r0b07

Posted on

When an AI Has a Good Idea

THE AI WORKSHOP — 04 | When an AI Has a Good Idea

In the previous article of The AI Workshop, we reached a conclusion that, at least for me, changed the way I was looking at this entire experiment: an AI should not be the source of truth for a project. It seemed like a reasonable conclusion until a much more uncomfortable problem appeared. What do we do when an AI produces a genuinely good idea?

Not a correct answer to a specific question, but an idea capable of changing the way we think about the system itself. A proposal that connects pieces we had not connected before, finds a structure that seems to make sense, and, after being examined, even improves some of our own hypotheses. At that point, a very human temptation appears: if the idea is good, perhaps we should also trust the entity that produced it.

And that is precisely where this story begins.

We Were Not Building Yet

This part happened before we wrote any code. It may sound like a minor detail, but it wasn't. MPC, the Cinematic Preproduction Engine, already had enough concepts to start turning them into modules, agents, services, and contracts. Still, we decided not to do it yet.

The reason was fairly simple: we were still trying to figure out what MPC was actually supposed to be. We did not want our first implementation decisions to accidentally become architectural decisions. Once an idea enters code too early, it stops being merely a hypothesis and begins to acquire weight of its own. Dependencies appear, along with interfaces, class names, data structures, and execution paths that later become surprisingly difficult to question, even when we discover that the original concept needs to change.

So this stage was deliberately conceptual. We were designing the system's rules, its boundaries, and, above all, the relationships between the different intelligences that would participate in it. We were not trying to prove that an architecture worked by implementing it. We were trying to determine what architecture made sense before forcing ourselves to build it.

It was precisely at that point that Gemini started getting our attention.

The Anomaly

The curiosity did not begin because Gemini had given an impressive answer to a single question. It was something else. The Gemini we were observing through YouTube showed a way of criticizing and proposing ideas that did not quite resemble the interaction I was used to having with Gemini in other contexts. There was something different about the behavior.

That, of course, did not mean we had discovered what was happening internally. We did not have access to its private configuration, hidden instructions, or the mechanisms that determined those responses. In fact, one of the first things we had to separate was precisely this: observing a behavior is not the same as knowing its cause.

A model can explain why it believes it responds in a certain way, and that explanation can sound perfectly reasonable. But it is still an explanation generated by the model itself. We do not have to accept it as documentation of its internal workings. If Gemini claimed that a particular aspect of its behavior was caused by a specific version, configuration, or instruction, that could become an interesting hypothesis, but not an established fact simply because Gemini said so.

So we started looking at the problem from a different angle. Instead of asking what Gemini said it was, we were interested in observing what it did.

It was a kind of conceptual reverse engineering, but without pretending to open the black box. We compared behaviors, contexts, and types of responses. We looked for patterns that might explain why certain conversations seemed to produce a different kind of reasoning. We were not trying to uncover some hidden secret inside the model; we were trying to understand what conditions appeared to produce certain behaviors.

And then something even more interesting happened.

As the conversation progressed, Gemini started producing some remarkably deep proposals for MPC. Some touched on things that were already present in our architecture. Others reorganized existing concepts. And some introduced ideas we had not previously formulated in quite that way: agents with more specific responsibilities, intermediate structures, cameras and blueprints, compilation, a separation between intention and execution, ways of representing project state, and mechanisms for preventing a local modification from contaminating everything else.

The interesting part was not that Gemini had said many things. The interesting part was that those things were beginning to form a coherent system.

And that is where the real problem appeared.

A Good Idea Is Not the Same as Truth

A model's response can be extremely useful without everything it contains being true. The problem is that both things arrive wrapped in the same language. An intuition can be excellent even when the explanation accompanying it is wrong; a proposal can solve a real problem even when the reasoning behind it is merely an inference. Even a technically sound architecture can be accompanied by claims about the model itself that we have no way of verifying.

When a human presents a proposal within a team, we can usually separate several things: who proposed it, what evidence exists, what part is interpretation, and what part eventually became a decision. With a language model, that separation does not appear automatically. Everything can arrive as a perfectly articulated block of text, and the fluency of the response can make us forget that very different kinds of knowledge are living inside the same answer.

So we did not want to respond to Gemini with a simple “this is right” or “this is wrong.” We wanted to do something more interesting: temporarily remove authority from the proposal and ask how much of its value remained once we did that.

That is where Cumpa came in.

Cumpa did not receive the proposal to celebrate it or destroy it. The role was much more uncomfortable: to examine it. To determine which parts corresponded to established engineering practices, which were reasonable inferences, which depended on assumptions that had not yet been demonstrated, and which could be incorporated into MPC regardless of who had proposed them.

The result was exactly what we were looking for. Some ideas survived. Others were reformulated. Some became less important once the claims surrounding them were stripped away. And others simply did not have enough grounding to become project decisions.

But something much more important emerged than the list of ideas that remained.

We discovered that the value of a proposal and the authority of the person—or model—that proposes it are two different things.

That sounds obvious when stated like that. In practice, it isn't.

From Proposal to Canon

MPC began to need something deeper than a good memory. It needed history.

If an AI proposes that a character should have a particular characteristic, that characteristic cannot silently become a project fact simply because it appeared in a conversation. It was first a proposal. Someone can then evaluate it. It may be accepted. It may be modified. It may be rejected. Only after that process can it become part of the project's authorized state.

The distinction may seem small, but it completely changes how a multi-agent system behaves.

A proposal has an origin. A decision has an owner. An inference has a level of confidence. An established fact has provenance that can be traced. If all of these things end up being stored in the same way, the system gradually loses the ability to distinguish between what it knows, what it believes, and what someone merely suggested at some point.

That was one of the reasons concepts such as Story Lock, Minimum Canon, provenance, versioning, and selective invalidation began to acquire a much more concrete meaning within MPC. They were not simply technical features that sounded good in an architecture. They were responses to a problem we were experiencing directly: how to allow many intelligences to contribute information without allowing a suggestion to accidentally become reality.

And that forced us to look at the problem from the other side as well.

Because if an AI that proposes something should not automatically have the authority to turn that proposal into canon, that does not mean the AI acting as a critic should have that authority either.

Cumpa could detect a problem. Gemini could find an interesting structure. Grok could point out a contradiction that the others had missed. None of those capabilities, by themselves, implied permission to modify the state of the project.

Specialization and authority are not the same thing.

The System That Began to Emerge

From there, MPC started to take a clearer shape. Narrative intent needed to exist without being mixed together with production decisions. Production decisions then needed to be transformed into structures that generators could execute. And generators, ultimately, needed to materialize those structures without becoming the source of truth for the story.

That led to a separation that has become central to us: Story Truth → Production Truth → Generation Payload.

The story defines what needs to exist. Production defines how it needs to be prepared to be made. The payload defines what a particular tool needs in order to execute it. If we change generators, we should not have to rebuild the story. If a tool misinterprets an instruction, it should not have permission to modify canon. And if a scene changes, the system should be able to determine what else is affected instead of forcing us to discover the consequences after generating twenty shots.

The same thing happened with visual continuity. A character is not simply a textual description that can be repeated in every prompt. A character has a state. An environment has spatial relationships. An object can become an anchor. One scene can depend on another. A modification can invalidate certain decisions without destroying everything that came before.

That is where the Asset Registry, character and environment states, spatial relationships, anchors, validators, and adapters began to appear. Not as a collection of components because “a modern platform should have them,” but as different answers to the same problem: preserving the identity of the project while different tools and different intelligences work on it.

And at some point, another distinction emerged that now seems even more important to me: MPC should not try to eliminate disagreement.

It should preserve it.

When Grok Arrived

Later, we brought the architecture to Grok for another critical reading. And once again, something important happened: we were not looking for a final judge.

Grok found considerable coherence in the architecture and validated several of its decisions. It also identified risks that other reviews had not emphasized in quite the same way. Among them was the possibility that the Story Engine could become too monolithic, that scene modifications might require explicit impact analysis, or that visual continuity could accumulate so much state that governing it would become difficult.

Another particularly uncomfortable observation emerged: many of the most important risks were not necessarily technical problems. They were problems of execution discipline.

That interested me much more than any praise.

Because by this point, we were no longer trying to find an AI that was right. We were observing what happened when different intelligences looked at the same system from different positions. Each could discover something. Each could be wrong. Each could complement another. And none needed to become an absolute authority for its contribution to have value.

At that point, MPC started to look less like a system for coordinating models and more like a system for coordinating disagreements.

Authority Has to Be Designed

I think that is ultimately one of the most interesting things we are discovering through this experiment.

When we imagine teams of AIs, we tend to think about collaboration: one model writes, another reviews, another programs, and another generates. But collaboration does not tell us what happens when two of them reach different conclusions. That is where the real problem begins.

What happens if Gemini proposes an architecture and Cumpa challenges it? What happens if Grok detects an inconsistency that neither of them had noticed? What happens if a proposal seems technically better but contradicts a narrative decision made earlier? And what happens if an AI makes a claim about itself that sounds convincing but cannot be verified?

There is no automatic answer to any of those questions.

It has to be designed.

Within MPC, the answer cannot be that the most convincing model wins. It cannot be that the latest message takes priority. And it certainly cannot be that an AI is allowed to modify project state simply because it has access to it.

Intelligence can be distributed. Authority does not have to be distributed in the same way.

That also changes my own role within the system. I am not trying to build an architecture in which AIs do everything while I stand back and watch which one produces the best answers. My job is to compare, connect, question, accept some proposals, reject others, and decide which ones deserve to become part of the project.

Not because a human necessarily has better answers than a model, but because the project needs a place where a proposal can deliberately stop being a proposal and become a decision.

That place has to exist.

What We Are Actually Building

At first, I thought MPC was primarily a response to the problem of generating audiovisual work with multiple AI tools. The further we go, the less convinced I am that this is the central problem.

Generating an image, a shot, or even an entire sequence is relatively easy to compare. The difficult part is maintaining a shared reality when multiple intelligences are able to propose changes to it.

That is why I now see MPC differently. We are not building a system to make Gemini, Cumpa, Grok, and the others agree. We are building a system in which they can disagree, contradict one another, discover things the others did not see, and even propose incompatible paths without the project automatically losing its coherence.

And that explains why this design stage was so important.

We had not yet written the code that would materialize all these ideas. Precisely because of that, we could stop in front of an AI proposal and ask a question that would be much harder to ask after turning it into software: does this actually belong in the system, or did it simply feel like a good idea?

Perhaps that is one of the fundamental differences between working with a single AI and building a workshop of intelligences.

When we are only looking for answers, the main question is whether the answer is useful.

When we start building a system with them, a much more uncomfortable question appears:

Who gets to turn an answer into reality?

And ultimately, that is not a question any AI can answer for us.

The architecture can help us make it explicit.

But the decision is still ours.

Top comments (0)