The two bets, and why this part tests neither
I came into this series with two hypotheses. First, that governed memory would outperform naive memory in an emergent communication task. Second, that optimization would rediscover a typed representation I had designed by hand for an earlier project. Both are interesting. Both need a working baseline to measure against.
So before I could test either, I had to answer a more basic question. Can two tiny neural networks, starting with no shared language, actually invent one that works? If the answer is no, then every clever architectural idea downstream is solving a problem that does not exist.
The only honest first question was not "is this profound" but "is there anything here at all." So I stripped everything out. No governance. No typed representations. No prior structure. Just the minimum machinery for communication to happen, run enough times to trust the result.
The game
I used a Lewis referential game. It is small enough to run on a CPU, which lets me run many seeds instead of one lucky one. The world is 256 objects. Each object is a bundle of 4 attributes with 4 values each, so every object is unique but they share underlying structure.
There are two agents, a sender and a receiver, with no shared vocabulary to start. The sender sees one target object and emits a fixed-length message of 2 symbols, drawn from a vocabulary of 16. The symbols are never seeded with meaning. Which symbol means what is something the agents have to invent.
The receiver sees the message and a set of 5 candidates, one target and four random distractors, and points at the target. Guessing at random gets it right 0.20 of the time. The sender learns by REINFORCE, updating on the reward when the receiver guesses right. The receiver learns by supervised cross-entropy. That split is standard for these games: the receiver learns to read signals while the sender learns to produce them.
Why a go/no-go, answer-blind, many seeds
Emergent communication experiments are notorious for looking good on a single run and falling apart across seeds. Agents can settle into a degenerate code that leans on positional cues, or converge on something that looks like communication but is really memorization. One pretty run proves nothing.
So I set pass thresholds before running anything, ran 8 seeds for 3000 steps each, and did not look until it was done. If the effect was not well above chance and consistent across seeds, the project was dead and there was no point going further. This was not a hunt for a good-looking result. It was a check that the substrate is real.
What happened
A decisive GO. On seen objects, accuracy was 0.970, far above the 0.20 baseline, so the agents clearly learned to map messages to objects. On its own that could be memorization, so I tested on held-out objects the agents never saw in training. Accuracy there was 0.937. That is the crucial number: they generalize over attributes rather than memorizing object identities.
Two controls rule out the boring explanations. Shuffle the messages and accuracy drops to 0.205, chance level, so the receiver genuinely depends on the message. Make the receiver message-blind and accuracy is 0.203, chance again, so there is no positional or candidate leakage doing the work. The agents also used 93 distinct messages, so this is not a one-symbol shortcut.
One number is quieter. Topographic similarity, which measures whether similar objects get similar messages, sits near 0.2. That is low.
The catch: capacity equals object count
The win here is narrow. A substrate exists. It is not a win for compositionality, and that low topographic similarity is the tell.
The channel capacity is 256, which is 16 squared, the number of possible two-symbol messages. That equals the number of objects. So the agents could in principle hand out one unique message per object and never learn a shared rule at all. Nothing in the task punishes that. The agents could have just invented 256 names. Proving they did more than that is the whole game.
The numbers say they landed in between. Held-out accuracy of 0.937 rules out pure memorized labels, because you cannot name an object you have never seen, so real attribute structure is leaking into the code. But topographic similarity near 0.2 says that structure is weak. This is mostly naming with a little systematic coding on top, not a compositional language. Compositionality had no reason to emerge, because naming was always available and always enough.
To force the question, I have to take the easy option away. What if I squeeze the channel, so there are far fewer possible messages than objects and a one-name-per-object code becomes physically impossible? Then naming cannot save them, and I get to see whether they can build something more structured instead.
Part 2 begins there.
Top comments (0)