Getting an AI agent to take an action is the easy part.
Getting it to know whether the action actually worked — that’s the problem nobody talks about until something goes wrong.
When I built the first multi-step actions for StareBrain (Google search, messaging, email — all running on-device via Android’s accessibility layer), I hit the same failure every time: the agent would complete its steps, report success, and be wrong.
Not maliciously. Just confidently wrong.
The root problem
An agent dispatches an action — say, sending a message. It checks: did I complete the steps? Yes. So it reports: SUCCESS.
But “I completed the steps” is not the same as “the message was sent.” The steps could have run against a stale UI state. The send button could have been intercepted. The accessibility event could have fired without the action completing.
The agent’s self-report is not ground truth. System state is.
What verification actually means
Before dispatching, define what success looks like:
`val SUCCESS_THRESHOLD = mapOf(
"stateChanged" to true,
"timestampAfterDispatch" to true,
"matchesExpectedOutput" to true
)`
After dispatch, check state — not the agent’s report of state:
`enum class ActionStatus {
CONFIRMED, // State verified against pre-defined threshold
FAILED, // State verified as unchanged
DENIED_UNRESOLVED // State unverifiable — not SUCCESS by default
}`
DENIED_UNRESOLVED is the important one. It’s not a placeholder. It’s a permanent first-class status for actions where the outcome can’t be verified. It does not decay into SUCCESS on timeout.
Concrete example: Messaging
After the agent attempts to send a message:
Check the sent folder for a message matching the expected recipient + content
Confirm the timestamp falls after dispatch
If both pass → CONFIRMED
If sent folder is unchanged → FAILED
If the sent folder is inaccessible or ambiguous → DENIED_UNRESOLVED
The agent never self-certifies. The artifact (message ID, sent record) is the proof.
Why this matters more for on-device agents
Cloud agents fail quietly — the error lives in a log somewhere. On-device agents fail visibly, in the user’s own apps. A message sent to the wrong person, an email drafted but not sent, a search that looped — the user sees it immediately.
That visibility is actually an advantage if you build for it. The verification layer I described above catches most failures before the user does. The ones it can’t catch get flagged honestly as DENIED_UNRESOLVED rather than silently logged as success.
Transparent failure beats confident wrongness every time.
StareBrain is an on-device AI agent for Android. Pre-launch — Waitlist here.
Top comments (0)