Haven't played Slay the Spire? Here's the background - the rules, and why the game is hard to learn.
Besides invalidating a lot of work, the resurgence of the block bug had a positive outcome: it reminded me of the basics of engineering. Anything you don't test is broken. This isn't superstition or Murphy's law: from my experience it's simply fact, driven home by many painful lessons during my career.
So, I finally stopped to consider the situation: I had two implementations of the game. One was the ground truth, the game played as intended, with graphics. The other was the game the agent played, running in headless mode through a series of untested patches. If my agent is to ever learn to play the former, the latter has to match it exactly. Thus began the quest for fidelity. I'd instructed Claude to set up a test harness. The game would be played using the same seed in both environments. The game's state would be compared after each step. Any divergence would be root-caused, fixed, and most importantly, would re-trigger testing of all previously-passing seeds.
Eight days after setting out on this laborious path, about 100 bugs were found and fixed, and I had 15 seeds running with perfect parity. I'll list some of the more interesting ones below just to give an idea of just how broken a game the agent had been playing.
Bug 11 was a big one. Setting up the game's combat state also involved enabling the "end turn" button. Because we had no such button, that line threw an exception and combat prep was aborted. A part of combat setup involved clearing powers (persistent effects that are meant to last during the combat), as well as resetting energy. It turned out our agent fought a hyperactive cultist, one that stacked their ritual more and more the longer training went on.
Similarly, bugs 14 and 15 had to do with energy relics. These powerful relics (usually received from beating a boss) increase the amount of energy a player starts their turn with. That amount was never decreased. When the agent got better and started beating bosses, it became easier and easier to do so, since subsequent runs enjoyed the effects of all previously-gotten energy relics.
If you're wondering how the agent suddenly started beating bosses, it may have been in part due to the fact bug 16 disabled all monster "change states". That meant, among other things, Lagavulin never woke up and the Guardian never assumed its defensive position.
Not all bugs were beneficial to the agent, of course. Similar logic bugs caused many relics to not do anything: bottling a card, gaining max HP from fruit relics and Potion Belt all originally did nothing, as did Maw Bank, Meal Ticket and Dream Catcher.
A couple of sneaky heisenbugs were artifacts of the harness, more or less. Slay the Spire uses a seed to generate a sequence of random numbers. The reason a run is reproducible is that given the same seed, the same numbers will be generated. One bug was run-of-the-mill uninitialized fields due to skipping a constructor. Another had to do with the card Discovery, where some bug in the game engine itself wasted one random number if a card was picked. All of these quirks had to be emulated in headless mode just to allow testing and fixing to be deterministic.
Another class of bugs was around events. Because events often have a unique GUI, and their buttons often have flavorful text, the agent would get stuck in them, or skip them illegally (to both benefit and detriment), or the event would crash silently. For example, the Gremlin Wheel's event labels have "prize!" and "prize?" for positive and negative outcomes. It's boring but some piece of code needs to click one or the other.
Gremlin Wheel was also representative of another class of bugs, around real time. The way it works is the agent clicks on "play", then the game plays an animation, silently counts Mississippis to let a couple of seconds of real time pass, then presents the result. Headless had no animation loop, so "ticks" (the unit of time) never progressed and the result was never reached. Conversely, act 3 has an event that lets you skip all the way to the act boss. Presumably to prevent speedrunners from abusing this, this event is only offered if 800 seconds have elapsed from the run's beginning. Headless mode's time was always 0, so this was never true, causing a divergence where for the same seed, different events were encountered.
You'll notice I don't mention the fixes - there was no clever engineering involved. In most cases, the appropriate code paths were recreated, sometimes using reflection, other times by re-implementing on a case by case basis. Claude Code has been invaluable in this, since the amount of manual labor was staggering.
The point of this post isn't just "haha bugs are funny". Networks are adaptable by design, even humble million-parameter ones. This is both a blessing and a trap. The agent played a completely broken game, but one that shared some principles with Slay the Spire. When we moved to per-card scoring over mean pooling, the agent consistently won fights it was losing. When we started rewarding block immediately, the agent got better at preserving HP. Things were directionally right.
The magnitude was wrong, as was the actual game being played, but nothing in a training loop could point to that - loss decreased, metrics improved, the agent was clearly learning. That's why it took me a long time to figure out our headless implementation had more bugs than correct lines of code.
The other thing I don't really mention is how each of these impacted agent performance. This is in part because it wasn't tested after every bug and because bugs made the game both easier and harder so there was no reason for any directionality. When the dust settled, the agent went from an average floor of 12.0 and beating the act 1 boss 6.5% of the time to 16.7 and 36%. Most importantly, I could be reasonably certain results can be ascribed to actual agent performance in Slay the Spire, rather than whatever jumbled mess the agent had been playing until that point.
That allowed me to focus again on improving the agent's quality of play. Ironically, because we moved to a search approach which relied on a simulator, we got to have the fun of aligning it to the real game all over again. More on search next week.
Top comments (0)