I built a harness to see whether a decision model could get a World of Warcraft character to level 5. The most frustrating part was watching it get stuck on uncertainty.
I used Astra in Codex to build it. During development, Sage kept overindexing on indecision, so I removed "uncertain" options from menus.
That sounds like a tiny change. It made me look much harder at the interface between a model and the actions it is allowed to take.
Disclosure: I'm a co-founder of Levanto, which makes Sage. This article was prepared with AI assistance from my development notes and the public code.
The menu is part of the program
In this kind of harness, the model doesn't receive a blank page and unrestricted control. It gets a current image, some context and a list of available choices.
That means a menu can create its own failure mode. If one option avoids committing, the system can keep choosing it while the controller asks essentially the same question again.
My removal of uncertain options was a development change, not a controlled experiment. I haven't measured its isolated effect. The saved run traces represent the menus used at the time they were recorded.
What I would examine in another system is the transition after the choice:
- Does it obtain new evidence?
- Does it change the view or the scope of the question?
- Does it consume a retry budget?
- Can it return to the exact same state forever?
An option named "inspect" isn't useful just because it sounds cautious. It needs to change what the next decision can know.
Input sent and action worked need separate records
Here's an illustrative trace shape. This is pseudocode, not an excerpt from the implementation:
decision_record = {
"observation": frame_id,
"available_choices": option_ids,
"chosen": choice_id,
}
input_record = {
"decision": decision_id,
"attempted": True,
"completed": True,
}
outcome_record = {
"input": input_id,
"new_observation": next_frame_id,
"effect": "unknown",
}
A completed keypress does not establish the effect in the world. For example, sending a cast command cannot by itself prove damage or a kill.
Keeping that uncertainty in the outcome record also matters for recovery. An unknown previous outcome shouldn't automatically prevent every possible next action. The controller needs to ask what the current evidence permits.
Revalidate at the last useful moment
There is another gap between the frame sent to the model and the frame present when the answer returns.
The public implementation's decision cycle checks task and session validity, input generation and image age. Its dispatch validation also checks that the new frame actually is new and still comes from the expected source and geometry.
The pattern I take from this is:
observe → ask → receive choice → revalidate → attempt input → observe effect
The revalidation step can reject an answer that was reasonable for the earlier screenshot. That rejection should be visible in the trace, so a stalled system isn't automatically blamed on the model.
What I can claim from this experiment
The character did reach level 5. The campaign record covers 32 sessions and 16 harness fixes. There were eight recoveries while the runner was stopped and one operator-selected hunting-area change.
Only the final session, roughly 18 minutes from late level 4, ran without human gameplay input, code changes or restarts. I don't have an experiment showing what the same harness with a simpler policy would have done.
That limitation is useful when reading agent demos: the outcome and the amount of development needed to get there answer different questions.
I came away wanting better traces for indecision: not just how often a model abstained, but what happened after each abstention and whether the next question had any new information.
Top comments (0)