DEV Community

Cover image for My Agent Returned Success. The Browser State Said Otherwise
Raju Dandigam
Raju Dandigam

Posted on

My Agent Returned Success. The Browser State Said Otherwise

An agent clicks Checkout. The browser tool returns success. The run completes without an exception.

But the page is still the cart.

Which result should the test trust?

This is one of the most dangerous green states in agent development: the action succeeded, but the outcome did not. A browser driver can accept a click without causing navigation. A queue can acknowledge a message that a consumer later rejects. A database client can resolve before a downstream projection updates.

If your test ends at the tool's return value, you are testing the transport—not the task.

Success exists at three layers

Consider three claims:

  1. The tool function returned normally.
  2. The agent run completed normally.
  3. The expected state transition occurred.

The first two are execution evidence. The third is outcome evidence. All three are useful, but they are not interchangeable.

agent intent
    |
    v
tool call ---------> { status: "success" }
    |
    | independent read
    v
browser state -----> { page: "cart" }
                           |
                           v
                 observed outcome: failed
Enter fullscreen mode Exit fullscreen mode

That last read closes a gap that ordinary happy-path tracing can leave open.

Record the contradiction, not just the exception

While maintaining AgentInspect, I added a synthetic recipe for this failure. The important idea is portable: perform the action, then observe the world through a separate surface.

This example is verified against agent-inspect@6.17.6:

import { inspectRun, observeOutcome, step } from "agent-inspect";

const browserState = { page: "cart" };

await inspectRun(
  "checkout-agent",
  async () => {
    const before = { ...browserState };

    const action = await step.tool(
      "browser.clickCheckout",
      async () => {
        // Synthetic bug: the tool reports success,
        // but the page never changes.
        return { status: "success" as const };
      },
    );

    const after = { ...browserState };
    const transitioned =
      before.page === "cart" && after.page === "checkout";

    await observeOutcome("checkout-transition", {
      expectation: "Page changed from cart to checkout",
      status: transitioned ? "passed" : "failed",
      method: "snapshot",
      actual: {
        beforePage: before.page,
        afterPage: after.page,
      },
      evidence: {
        actionStatus: action.status,
      },
    });
  },
  {
    traceDir: ".agent-inspect",
    silent: true,
  },
);
Enter fullscreen mode Exit fullscreen mode

The run status is success. The observed outcome is failed. Both statements are true, and preserving both is far more useful than forcing them into one status.

Make the mismatch visible in review

Inspect only the observations section:

npx agent-inspect report <run-id> \
  --dir .agent-inspect \
  --section observations
Enter fullscreen mode Exit fullscreen mode
## Observed outcomes

Total: 1 (passed 0, failed 1, unknown 0, skipped 0)

| Name                | Status | Expectation                         | Method   |
| ------------------- | ------ | ----------------------------------- | -------- |
| checkout-transition | failed | Page changed from cart to checkout  | snapshot |
Enter fullscreen mode Exit fullscreen mode

Then make the same evidence fail CI:

npx agent-inspect check <run-id> \
  --dir .agent-inspect \
  --fail-on-observation failed
Enter fullscreen mode Exit fullscreen mode

The check exits nonzero even though the agent run completed. You can also find this class of run later:

npx agent-inspect search \
  --dir .agent-inspect \
  --observation failed
Enter fullscreen mode Exit fullscreen mode

This is the difference between recording an interesting fact and making that fact operational.

Choose an observer outside the action

The strongest observer reads a different surface from the one that performed the action.

Action Better observation
Browser click DOM, accessibility tree, or URL snapshot
Database write Independent read by business key
Queue publish Consumer receipt or downstream state
File generation File existence plus a content or digest check
API mutation Follow-up GET or emitted domain event

If clickCheckout() performs the click and verifies success using its own internal didClick flag, the test can repeat the same mistake twice. Independence matters more than sophistication.

For eventually consistent systems, the observation may need a bounded polling policy:

async function eventually<T>(
  read: () => Promise<T>,
  accept: (value: T) => boolean,
  attempts = 5,
): Promise<T> {
  let value = await read();

  for (let i = 1; i < attempts && !accept(value); i++) {
    await new Promise((resolve) => setTimeout(resolve, 200));
    value = await read();
  }

  return value;
}
Enter fullscreen mode Exit fullscreen mode

The timeout, interval, and accepted state should come from the product's consistency contract—not from whatever delay makes today's test pass.

Define “done” before you instrument it

Outcome checks become much sharper when the team writes the completion condition before choosing a selector or API call.

For checkout, “the button was clicked” is an implementation detail. “The browser reached the checkout route and displayed the expected order context” is closer to a product outcome. For a support agent, “the ticket API returned 201” may still be weaker than “the created ticket is visible by its business key with the expected queue and priority.”

A small outcome contract keeps that distinction reviewable:

type OutcomeContract<T> = {
  name: string;
  expectation: string;
  observe: () => Promise<T>;
  passed: (value: T) => boolean;
};
Enter fullscreen mode Exit fullscreen mode

The contract does not need to be a framework. Its value is forcing the action and the proof of completion to be named separately.

Keep the evidence smaller than the incident

actual and evidence should contain the smallest proof that explains the result: status, field presence, counts, digests, or safe identifiers.

For this case, beforePage, afterPage, and actionStatus are sufficient. A full HTML document, screenshot sequence, customer record, prompt, or browser session is not.

AgentInspect bounds and redacts recorded values before disk, but that is not permission to capture everything. Collection policy comes first; redaction is defense in depth.

What this pattern does not prove

An observation is only as trustworthy as its observer. A stale replica may report failure after a successful write. A DOM selector may target the wrong element. A passed observation also does not prove the whole workflow was safe.

This is post-execution evidence, not a runtime guardrail. Permissions, idempotency, budgets, and destructive-action controls still belong in the application and tool gateway.

Try the failure without a browser or model key

The pinned browser/MCP observed-outcome recipe uses only in-memory state. It deliberately returns tool success while leaving the simulated page on cart.

That makes it a useful five-minute exercise: run it, inspect the failed observation, then replace the synthetic snapshot with one state read from your own system.

The conceptual change is small but consequential: stop asking only whether the agent ran. Ask which independent state would prove that the work happened.

What surface would you trust after your agent says done?

Top comments (2)

Collapse
 
raknaos profile image
Baptiste Le Bouquin

This maps directly to a failure we hit in production with headless agents driving real logged-in sessions: an agent clicked a submit button, the tool returned success, the run reported completion — and the page was still showing the form, because a cookie-consent overlay had been sitting on top and swallowed the click. Everything up to and including the CDP response was green; only the DOM disagreed. Tool-level success really is transport evidence, nothing more.

What worked for us is essentially your separate-surface idea: after every state-changing action, a fresh read-back of the things that only exist if the outcome happened — URL, heading text, presence of the confirmation element. We treat the read-back as the only evidence worth recording. Optimistic UI makes it trickier: the confirmation element can exist for a couple of seconds and vanish once the server rejects, so we learned to sample past the optimistic window instead of reading immediately.

How does observeOutcome handle the timing side? The naive version — read state right after the action resolves — will disagree with reality whenever the app commits asynchronously. Did you end up with a settle heuristic, or does the test author declare what to await?

Collapse
 
kevinpruett023_kevinpruet profile image
kevinpruett023 kevinpruett023

Hi. good post. Let me know which chatting platform you use. Let's discuss further about collaboration. how about you?