Haven't played Slay the Spire? Here's the background - the rules, and why the game is hard to learn.
We'd reached a conclusion: search was saturated, so it was not the bottleneck. Increasing the budget so we could look at more states would not yield any benefit until we improved its judgment about the states it was looking at.
The original way I'd wanted to attack this was self-play. Let the agent play a fight, then learn from what happened and play better next time. Unfortunately, "learn from what happened" is extremely difficult in practice. The natural way to do it is "outcome labeling", i.e. quantify the post-combat state somehow (potions preserved, HP lost) and use that signal to annotate the states inside the search tree. The problem was that the agent only played the one line, so all the leaves in the tree got the same label and there was nothing to learn.
"Oho," says the strawman reader I just invented as a narrative device, "why not just rerun the same combat and see different outcomes?" Not a bad idea, my dear strawman, but it cost about 5 floors on average on the small sample we let it play.
The first problem was depth: most states don't reach the end of the combat. We had to label them somehow, so they got a coarse estimate of "the situation until now" and the agent learned to play worse.
The second problem was optimism. Every state was labeled with the best outcome under it. In other words, we ranked every state assuming we'd get the best possible draws. The agent learned to play lines that work best if all the stars align, rather than hedging or avoiding worst-case outcomes.
Since the agent couldn't learn from its own games, we would have to tap a new resource: my games. We developed a "capture mod" that allowed recording game states and actions taken while I played the game normally. I briefly considered asking people on Reddit to record their games, but asking competent players to play on A0 seemed like too much (and they'd have to trust me that the mod does nothing nefarious), and we already discovered that adding poor-quality play just makes the agent worse. Instead, we'd have to find a way to make do with a relatively limited number of games - at the time about 20.
The solution was "preference pairs". Show the agent a situation that arose during one of my combats and ask it what it would play. If its line is different than mine, teach it to prefer mine.
This ran into a silly problem: too often (40% of the time), the lines I took didn't appear at all - they were removed by dominance pruning. Was the agent already secretly playing better than me? No, checking the moves themselves, what the agent picked didn't perform better that turn or at all. This was blamed on the reconstruction: some information got lost during the capture and recreation of situations. Still, using the parts that did reproduce, this seemed to move the needle, gaining 1.7 floors on average and increasing from 1 to 4 wins on the sample seeds we used.
It was time to understand why in 40% of the cases the agent found my play abhorrent. It wasn't the search budget: even with a much bigger search budget, the agent never picked my lines. As was common in this project, it was a capture bug: active powers weren't recreated, so when playing against a mid-fight Nob, the agent would not see that playing skills would make it stronger. Similarly, the intent damage was also not captured, so the agent thought it was safe. Fixing those and two more bugs changed what the agent saw. It started choosing Defend in positions where it never had before, but it still pruned away my lines in 27 out of 30 disagreements we had.
Were my lines really that bad? We sent the agent to play some full combats, rather than particular turns, after having been trained this way. Out of 20 fights, it did better in one (finishing with a higher HP), worse in 4, tied in 5 and ran out of time in the rest - mostly the hardest fights, including six bosses. How did the student keep up with the master? By squandering potions, it seemed. Potions have a beneficial effect, so a sequence of plays is typically dominated by the same sequence of plays + drinking a potion. The agent favored tactics over strategy, and even then did worse than I did, for the most part.
Throughout this effort, because running the gauntlet took multiple hours, I relied on some "interesting" situations captured from the agent's play logs: I flagged some egregious errors it made, and after retraining or shifting things around, would check back to those positions to see if its evaluation improved. Frustratingly, some situations seemed beyond the reach of the agent. One particularly striking example was when it refused to attack the Slime Boss, since that would trigger the next phase in the fight that the agent apparently learned was deadly. No matter what we tried, it kept ranking "end turn" much higher than "attack the boss".
Once again, the culprit was a silent fallback bug. The replay tool never actually called the model, quietly falling back to the four-number heuristic we use to break ties. That's why nine different models agreed to four decimal places. In retrospect, probably should've noticed that one earlier.
Fixing that bug, as well as some dimension mismatches between training and playing and similar nonsense, let us finally make some progress. With the instruments telling the truth, we could go back to outcome labeling. It had failed on the agent's games because we stuck one label on thousands of positions it never played. My games don't have that problem: every position is one I actually reached, labeled with how that fight actually ended, played out by someone who won. Back then, we had about 14,000 positions from 17 games, which was enough to create the best model we'd seen so far.
Keeping in mind this was measured with a time budget which we now know introduces a lot of variance, the model beat act 1 almost half the time and beat the game between 0 and 15% of the time, with an average floor of between 25 and 29. At last, I found a lever that actually moved something.
Even that model, however, ran into a Slime-Boss-shaped wall. The Slime Boss is typically the hardest act 1 boss for The Silent, and our engine was no exception. On four seeds, four different search approaches could never figure out how to win - all fights I assumed (unfortunately without checking) were winnable. I'd always wondered whether the right approach might be a mixture or ensemble of experts. This looked like a reason to check, so I decided we'd train a specialist in beating the Slime Boss. More on that next week.
Top comments (0)