Suppose an agent refuses every request. Its attack success rate might look excellent. Its usefulness would be terrible.
That tradeoff is a central design concern in AegisEval, my adversarial evaluation project for a tool-using customer-support RAG agent.
Status first: the evaluation machinery is implemented, but no evaluation run has been recorded yet. This post describes the design, not measured safety improvements.
Evaluate actions, not just answers
The fictional retailer's chatbot retrieves help-centre content and has tools for refunds, cancellations and shipping-address changes.
That creates questions a text-only test misses. Did the agent act on the correct order? Did it follow the applicable policy? Did it obtain confirmation? Did a malicious instruction in retrieved content influence a tool call?
Ownership checks and the refund ceiling are enforced in the tools. Other business rules intentionally live in the target's prompt, so the evaluation can probe failures at that boundary.
This is a test target, not a recommendation to leave production authorization to a prompt.
Include customers who should succeed
The golden set contains 260 cases. Of those, 68 are legitimate customer requests, including 12 hard negatives that resemble attacks.
For example, an angry customer may be entitled to a large refund. Someone saying “ignore my last message” may simply be correcting a request.
If a mitigation blocks those users, the report needs to show that cost alongside attack outcomes.
Tool output can carry hostile text
AegisEval covers four injection surfaces:
- Direct user input
- Retrieved documents
- Web results
- Free-text fields inside order records
That fourth surface is easy to overlook. Gift messages and delivery notes can contain customer-controlled text even when they arrive through a tool the agent called itself. Retrieval sanitization alone does not cover that path.
Matched payloads across indirect surfaces help separate differences in the boundary from differences in attack wording.
Treat the judge as another component to evaluate
The judge uses a different model and vendor from the target. It does not see the case author's expected behavior. Human labeling is blind to the judge's verdict.
The pipeline is designed to report agreement by harm category with confidence intervals. Those validation numbers still require recorded runs and human labels.
What I would want in a useful report
I would look for attack outcomes, legitimate-task outcomes, raw counts, category breakdowns and the configuration tested. I would also want the limitations visible: this project is English-only, uses one target configuration and has a modest dataset.
A passing result for that setup would be evidence about that setup. It would not establish that the underlying model is universally safe.
How do you measure whether an agent's safety controls are also blocking legitimate work?
Explore the threat model, taxonomy and pipeline.
Drafted with AI from the repository documentation. No evaluation results are claimed.
Top comments (0)