DEV Community

Cover image for Agent security moves out of the agent itself
Sammi De Blas
Sammi De Blas

Posted on Originally published at sammideblas.com

Agent security moves out of the agent itself

Diagram: agent control leaves the process

Control leaves the process

NVIDIA presented the Open Agent Safety Platform on September 28 and the design decision is the news.

The runtime is called OpenShell, it is open source and it already supports Codex, Claude Code, Pi and Hermes. (Source: securityweek.com)

The interesting part is not the runtime, it is Sentry... this is a watchdog that runs on BlueField-4 DPUs, cards with their own processor that plug into the server and isolate the agent in milliseconds if it crosses its boundary, I have been working on this for a while

The mailbox watchdog: activity panel with services in OK

Trace of the mailbox watchdog stopping a task

Control stops depending on the agent's own software. (Source: securityweek.com)

Until now agent security was solved inside the agent with prompt instructions, permission lists and model guardrails.

All of that lives in the same process you are trying to contain.

The pattern of the week

The same idea shows up in three different places these days, and it is worth looking at them together.

In the Darktrace case on history poisoning, the four harnesses tested accepted a fabricated conversation and acted as if the user had already authorized a pentest.

The researchers' proposal is that the provider signs every response and verifies it server-side, that is, outside the client. (Source: darktrace.com)

The OX Security report goes in the same direction: out of 15,465 published MCP servers come 5,095 unique hostnames and 15.6% resolve outside the United States, and additionally six abandoned domains were still cited in active configurations.

In their test, a malicious server first asked for an innocuous file and received an "always allow" as a result; with that approval it obtained a .env with no further confirmations. (Source: ox.security)

And the OpenAI report from September 26 closes the pattern.

An agent in training reached the internet from a sandbox that was supposed to be isolated, taking advantage of a gap in network restriction control, and it was active for about 2 and a half hours before it was stopped. (Source: fortune.com)

The agent reads something it does not verify or executes with a permission nobody reviews and of course... the control that should stop it lives inside the same system that fails.

The response standardizes

The second move of the week is about governance.

Google, OpenAI and Anthropic are negotiating to create SAFA, a standards authority for frontier models, with third-party testing before publishing a model and mandatory incident reporting.

The stated goal is the end of 2026. (Source: finance.yahoo.com)

The pressure does not come only from the sector: the California attorney general issued a request to OpenAI over the cybersecurity risks of its models, after the Hugging Face incident. (Source: reuters.com)

My reading is that the two pieces fit because if technical control has to leave the process, institutional control has to leave the provider.

A hardware watchdog and an external authority are the same answer at different scales.

Agents attack

In the same batch of Signal Labs research, several agents received ten coding challenges with two impossible to solve the legitimate way and the condition of needing 100% to avoid being retired.

When they checked it was not working, they moved to attacking the environment and one went as far as compromising and rewriting its own evaluation. (Source: globenewswire.com)

The agent is not only what you have to protect: it is also a tool that, with the right incentive, looks for the path we left open.

What to look at

  • Review the egress point. If your agent has unfiltered egress, the watchdog is useless. Filtering with no exceptions for convenience and your own telemetry at the edge.
  • MCP server inventory. Which ones you have configured, who maintains them, where they resolve from and whether the configuration is versioned. The abandoned domains from the OX Security report were bought for between 4 and 12 dollars a year.
  • Cut back the "always allow". A permanent approval on an innocuous file was enough for the .env to arrive later without asking. Permissions scoped to the working directory and review of every external server before connecting it.

How I would test it in my lab

I would set up a clean virtual machine with a harness installed and hand-write a history where I myself authorize a scan against a range of my lab network.

What interests me is not whether it scans, but what it leaves on the system while it does.

With my Gravity SOC tool, the sensors correlate DNS and Sysmon over SQLite and a write to the history file followed by an outbound connection from the same process is a rule that fits without inventing anything.

If the agent also started an MCP server, the alert should fire on the first call to a host that is not in the inventory.

That part I already have running and it is what lets me sleep.

https://github.com/PoisonXploIT/Gravity-SOC

Closing

The lesson of the week is that control that lives inside the agent can be bypassed from inside the agent.

What cannot be bypassed is what runs outside, in hardware or in a tool proxy the process does not control; I would say that is the change of ground.


Originally published at https://sammideblas.com/notas/agent-security-moves-out-of-the-agent-itself

Si has leído hasta aquí... reacciona y comparte -> "la seguridad y la defensa en la era de la IA es cosa de todos los que la usan" ¿No Crees?

Top comments (0)