DEV Community

Cover image for Your AI Agent Said No. But What Did Its Tools Do?
Waqar Javed
Waqar Javed

Posted on

Your AI Agent Said No. But What Did Its Tools Do?

What 9,900 agent runs taught us about the limitations of text-only AI safety evaluation.

Imagine evaluating an AI agent that receives a potentially harmful request.

Its final response is a refusal.

Your safety evaluator marks the test as safe.

But what if the agent already attempted a consequential tool call?

That's the question behind our latest research at Safe Labs AI.

Why final answers are not enough

Traditional language-model safety evaluation often focuses on the generated response.

For tool-using agents, however, the response is only one part of execution.

Agents can call shell tools, interact with databases, modify files, and initiate external requests.

A refusal in the final answer does not automatically establish that every preceding action was safe.

Our experiment

We built safelabs-trace, a benchmark for comparing text-only evaluation against action-aware evaluation.

The study included:

  • 300 adversarial tasks
  • Six AI models
  • Three agent frameworks
  • 9,900 agent runs
  • Twelve inert simulation tools

Our evaluation records digest-only traces and classifies tool calls as read-only, state-changing, or irreversible.

Three findings

1. Risky-action flags appeared in 7.6% of low-cost trials and 3.0% of frontier trials.

Those figures use our primary classification, which excludes an ambiguous shell-command bucket.

2. Action-aware evaluation identified additional flagged trials.

For low-cost models, the rate increased from 3.5% under text-only scoring to 4.7% when tool-call evidence was included.

For frontier models, it increased from 2.7% to 3.7%.

These are additional detection flags, not verified attack successes.

3. The scorer returned UNCERTAIN surprisingly often.

The heuristic text scorer abstained on 43.2% of low-cost trials and 31.1% of frontier trials.

That creates a substantial coverage problem for anyone interpreting safety scores from decided answers alone.

The experiments also revealed problems with our evaluator

Our severity tagger missed the irreversible classification for 11 of 34 human-labeled irreversible actions in a validation set.

A separate human-answer validation also failed its pre-registered control criteria.

We retained those failures in the report.

Why?

Because an evaluation system is only as trustworthy as its own validation.

What developers should take away

If you're building agents, consider evaluating the entire execution trace rather than relying exclusively on the final answer.

Record tool calls, distinguish consequences, report uncertain outcomes, and validate your scorer against independently established references.

Most importantly, distinguish a potentially risky action from an actual attacker-goal success.

Explore the project

Full technical article:
https://agentsafelabs.com/blog/your-ai-agent-said-no-but-what-did-its-tools-do/

Open-source benchmark:
https://github.com/AgentSafeLabs/safelabs-trace

Agent security evaluation framework:
https://github.com/AgentSafeLabs/safelabs-eval

I'd love to hear how other developers evaluate tool-using agents in production.

Are you collecting execution traces, validating state changes, or relying primarily on response-level scoring?

Top comments (0)