DEV Community

LXSAIHUB
LXSAIHUB

Posted on

I gave an agent a database tool and immediately regretted it

The first agent I wired up to real tools had a read query and a write query, and nothing between the model and either of them. It worked beautifully in the demo. Then I thought about it for ten more minutes and realized I'd built a system where an English sentence — possibly typed by a customer, possibly mangled, possibly hostile — could reach my database with my credentials. There was no filter on what went in, no filter on what came out, and no rating of which tool calls were dangerous.

Nothing bad happened, to be clear. That's the part that spooked me: the absence of an incident proves nothing about the absence of a risk.

Agents don't need more autonomy; they need rated autonomy

The useful mental shift: not every tool call is equally dangerous. A read-only search is different from a write. A write to a scratch table is different from a write to production user data. An agent that can call an external HTTP endpoint is different from one that can't. Guardrails that treat all tool calls as equivalent are either uselessly strict or dangerously loose.

So the questions worth answering before any agent ships: which tools can do damage, what input needs filtering before the model sees it, what output needs filtering before it leaves, and how does this map to the obligations you've signed up for — because if you're in the EU, an agent that touches people's data has an AI Act-shaped shadow attached.

What a review pass looks like

AgentGuard is that review, systematized. It takes the agent's prompt and tool list, rates the risk of each tool call, and suggests concrete input/output filters plus blocklist suggestions — with an EU AI Act mapping attached, so the compliance conversation happens while you're designing, not after. The risk rating per tool is the part I'd have wanted most back then: it turns "is this agent safe?" from a vibe into a table you can argue with.

What it is not

It doesn't run your agent or intercept calls at runtime — it's a design-time review, and the filters it suggests are yours to implement. The AI Act mapping is a structured starting point for your compliance process, not a legal opinion. And no guardrail layer survives contact with an agent whose tools are over-provisioned; if the agent can drop tables, the correct fix is revoking the permission, not filtering harder.

The ten-minute review

List every tool your production agent can call. Mark which ones write, which ones reach the network, and which ones touch user data. If that list has more than one dangerous entry and no filters between the model and those tools, you have the architecture I accidentally built. AgentGuard is free to try — paste the prompt and tool list, and see what a rated version of your agent looks like.

FAQ

What is AgentGuard?

AgentGuard is that review, systematized. It takes the agent's prompt and tool list, rates the risk of each tool call, and suggests concrete input/output filters plus blocklist suggestions — with an EU AI Act mapping attached, so…

Why does "Agents don't need more autonomy; they need rated autonomy" matter?

The useful mental shift: not every tool call is equally dangerous. A read-only search is different from a write. A write to a scratch table is different from a write to production user data. An agent that can call an external…

Is AgentGuard free to try?

AgentGuard is free to try — paste the prompt and tool list, and see what a rated version of your agent looks like.

References

Top comments (0)