DEV Community

Jules Robineau
Jules Robineau

Posted on Originally published at jrobineau.com

Your AI Assistant Can Reply to a Stranger. Do You Let It?

For months, an agent has been reading my inbox. An agent is a program driven by a language model, and it can act: read, sort, write. Mine handles mission proposals. It sorts them, evaluates them, and drafts the replies.

One question remains: who presses send?

TL;DR: my agent reads everything and drafts everything, without asking me. But for sending, each type of reply has its own rule. A simple question goes out alone, twice at most, with a delay. A refusal always waits for my click. A serious mission always comes back to me. And these rules live in code: no prompt can change them.

This article is for developers who want to wire an agent into their inbox without handing it their signature.

OpenClaw made the question concrete

OpenClaw is the open source AI assistant that blew up this year. It runs on your machine. It reads your messages. And it can answer them for you.

Security did not keep up. Censys counted 21,639 instances open on the internet. The extension marketplace hosted over a thousand malicious packages. Krebs on Security ran an investigation.

Those flaws get fixed with patches. But one question survives, even fully patched: do you let a machine talk to a stranger in your name?

I live with that question. My agent runs in production, on my real inbox. Here are the rules I gave it, and why.

The real risk is what leaves

When the agent reads, I risk my data. That is serious, but manageable: it runs locally, with minimal access.

When the agent writes to a stranger, I risk my name. A bad draft takes two seconds to fix. A bad email, sent to a prospect, cannot be erased.

So my rule is not about what the agent can do. It is about what leaves.

What the agent does on its own

All the reading and all the preparation run without me.

First, a dumb filter. The email must contain at least two words from the mission world. No AI here: a word list, fast and predictable. Newsletters stop there.

Then a local model reads the email. It fills a five-criteria grid: stack, contract type, start date, rate, remote policy. Each criterion gets a verdict: GO, NO_GO or unknown.

Finally, a cloud model drafts a reply in my style. The result: a draft. Nothing else.

I review none of these steps. If the agent gets something wrong here, it costs nothing. Nothing has left.

What it sends alone, and what it never will

When it is time to reply, the agent picks an action. Each action has its own sending rule. It is a Go switch, not a model mood.

The rate or remote policy is missing? The agent asks, and sends it alone. A simple question commits nothing.

Declining a mission? Never alone. A refusal closes a door, with my name at the bottom. The draft waits for my click. In the code, auto-send is hardcoded off for that action.

Sending my skills file? Never alone either.

And when a mission checks at least three criteria out of five, with zero NO_GO? The agent stops replying. It summarizes the thread and hands it back to me.

Look at the logic. The more serious the conversation gets, the fewer rights the agent has. Never the other way around.

Two guardrails, even on simple questions

Even the automatic sending of questions keeps two limits.

A cap: two automatic sends per conversation. It is a counter in code, not a promise in a prompt. On the third exchange, everything becomes a draft again. A loop with a stranger must stop on its own.

A delay: each send waits between two and six hours, at random. The reply looks human. And while it waits, the draft sits visible in my mailbox. I delete it, the send fails. That delay is my veto right.

The rules live in code, not in the prompt

My agent starts in a mode, picked by a flag: draft, hybrid or review. In draft, nothing ever leaves. In hybrid, only simple questions leave alone. In review, everything waits for my approval.

The prompt decides none of this. That is the most important point in this article.

Why? Because of prompt injection. A booby-trapped email can talk a model into changing its mind. It cannot change a compiled flag.

I already tested this philosophy on permissions: my AI agent tried to delete my secrets, it could not. Sending follows the same rule. It is enforced outside the model.

And when I want to approve everything: review mode

Sometimes I want zero automatic sending. I switch to review mode.

Every reply then lands in a small web interface, at home, never exposed to the internet. Three buttons: approve, edit, reject. If I reject, I explain in one sentence, and the agent rewrites the draft with my feedback. Every decision goes to a log, for retraining later.

For LinkedIn, I go even further. An extension pastes the reply into the message box. It never sends. The final click stays mine, physically.

The checklist for your agent

One rule per type of action, never per conversation.

  • [ ] List what your agent can do, action by action
  • [ ] For each action, three questions: reversible? my name? a stranger?
  • [ ] Let reading and preparation run without validation
  • [ ] Allow solo sending only for simple questions, with a cap
  • [ ] Add a delay before every auto-send: that is your veto right
  • [ ] Put the rules in code, a flag or a switch, never in the prompt
  • [ ] Keep a trace of every decision, to reread later
  • [ ] Keep the approval interface local, away from the internet

What to remember

So, do you let it? Reading, sorting, preparing: yes, fully. Talking to strangers: each action has its rule, with a cap, a delay, and a mode where everything goes through you.

OpenClaw did not create this problem. It scaled it. Your agent's freedom is an architecture decision, not a prompt line.

Want to wire an agent into your inbox without handing it your signature? Let's talk.


Sources: Krebs on Security, "How AI Assistants Are Moving the Security Goalposts" · Reco, "OpenClaw: The AI Agent Security Crisis" · Palo Alto Networks, on autonomous agent identity

Top comments (0)