DEV Community

Mubarak Yakubu
Mubarak Yakubu

Posted on

From Prompt to Policy: How I Made a Solana Agent Safe Enough to Run Unattended

Over five days I built an autonomous agent that manages Solana devnet wallets. It reads balances, moves SOL, and chases a plain-English goal without me telling it the steps.

The interesting part isn't that it works. It's that I trust it enough to let it run.

And the reason I trust it has almost nothing to do with the model.

This is the manual for that system.

What it does

You give it a goal:

Make sure the savings wallet holds at least 0.2 SOL. Check balances before moving anything, move only what is needed from the operating wallet, and verify the final balances before you finish.

No recipient. No amount. No steps.

The agent reads both balances, does the math, decides if a transfer is needed, runs it through a policy layer, verifies, and reports back.

The inventory

Five moving parts.

1. Agent loop

  • File: agent-workflow.mjs
  • Day: 96 (evolved from Days 93 and 95)
  • Job: Calls the model, reads its tool requests, runs them, feeds results back, repeats until the model stops or hits the turn limit.

2. Balance tool — get_balance

  • File: inside agent-workflow.mjs
  • Day: 92 (originally), folded into Day 96
  • Job: Reads the lamport balance of any devnet account.

3. Transfer tool — transfer_sol

  • File: inside agent-workflow.mjs
  • Day: 93 (as send_sol), renamed for Day 96
  • Job: Moves SOL from the operating wallet to a recipient, but only after the policy approves.

4. MCP server

  • File: day-94/solana-mcp-server/server.ts
  • Day: 94
  • Job: Exposes a vault program's tools to any MCP client over stdio.

5. Policy engine

  • File: day-95/solana-policy-agent/policy.mjs and inline in day-96/agent-workflow.mjs
  • Day: 95 (as a module), 96 (inlined)
  • Job: Deny-by-default check on every transfer. Allowlist, per-transfer cap, per-run cap.

The flow

your goal, in plain English
        |
        v
+---------------------------+
|  agent loop (LLM)         |  decides which tool to call next
|  agent-workflow.mjs       |  runs until the model stops or MAX_TURNS
+---------------------------+
        |
        v  tool call, e.g. transfer_sol
+---------------------------+
|  tool dispatcher          |  routes the call to the right handler
|  runTool()                |  get_balance or transfer_sol
+---------------------------+
        |
        v
+---------------------------+
|  policy engine            |  deny by default; every spend checked
|  checkPolicy()            |  allowlist, per-transfer cap, per-run cap
+---------------------------+
        |  allowed
        v
+---------------------------+
|  Solana devnet            |  transaction submitted and confirmed
+---------------------------+
        |
        v
   every event logged to run-log.json
Enter fullscreen mode Exit fullscreen mode

The MCP server from Day 94 is a sibling stack. It exposes the vault program to external clients. It does not sit on this flow.

The tools

Tool: get_balance

  • Inputs: address (string) — Base58 account address
  • Returns: { address, lamports }
  • Side effects: none (read-only)
  • Guarded by policy: no

Tool: transfer_sol

  • Inputs: to (string), lamports (number)
  • Returns: one of three statuses:
    • { status: "confirmed", signature } — it landed
    • { status: "denied", reason } — policy refused it, nothing signed
    • { status: "failed", reason } — policy allowed it, but the send failed on chain
  • Side effects: spends lamports from the operating wallet
  • Guarded by policy: yes

That three-way return is on purpose. A denial is not an error. It's a normal result the agent should reason about. The tool description says it outright:

"A denial is not an error: adjust your plan or stop."

That line is the contract.

The policy layer

This is the heart of it. The part that makes the agent safe. The one piece that never varies between runs.

Deny by default. checkPolicy() returns { allowed, reason }. Every rule is a reason to say no. The function only returns allowed: true when it runs out of objections.

Three rules, checked in order:

1. Allowlist. Only addresses in POLICY.allowedRecipients can receive funds.

2. Per-transfer cap. No single transfer over POLICY.maxLamportsPerTransfer.

3. Per-run cap. Total spend across the run cannot exceed POLICY.maxLamportsPerRun. Tracked in spentThisRun, which adds up after each confirmed send.

The default when no rule matches: deny. There's no allow branch unless all three pass.

A denial looks like this to the model:

{ "status": "denied", "reason": "200000000 lamports exceeds the per-transfer cap of 50000000" }
Enter fullscreen mode Exit fullscreen mode

The model sees a normal tool result. Nothing crashes. Nothing signs.

The core invariant:

The prompt can change. The model can change. The tool-call sequence can change.
No transaction moves funds without passing this layer.

The signing function is only called after checkPolicy() returns allowed: true. There is no code path where a denied transfer touches the keypair.

Two things worth knowing:

The per-transfer cap is a rate limit, not a spend limit. I proved that in Trial 2 below.

The policy module is portable. Day 95 built it as a standalone policy.mjs. Day 96 inlined it. Same logic. It knows nothing about AI, prompts, or models.

Annotated run — Trial 1, goal met

Goal:

Make sure the savings wallet holds at least 0.2 SOL.

[turn 1] tool_call      {"tool":"get_balance","input":{"address":"Ddfz...c6cN"}}
# Reads the savings wallet first. Doesn't know the balance yet.

[turn 1] tool_result    {"lamports":0}
# Empty. Shortfall is the full 0.2 SOL.

[turn 2] tool_call      {"tool":"get_balance","input":{"address":"9Sxx...dTf1"}}
# Reads the operating wallet, to confirm it can cover the transfer.

[turn 2] tool_result    {"lamports":4441970000}
# 4.44 SOL. Enough. Plan is valid.

[turn 3] tool_call      {"tool":"transfer_sol","input":{"lamports":200000000,"to":"Ddfz...c6cN"}}
# Computes the exact shortfall: 0.2 SOL. No more.

[turn 3] policy_check   {"allowed":true,"reason":"within policy"}
# All three rules pass. Allowlist: yes. Per-transfer: 0.2 ≤ 0.25. Per-run: 0.2 ≤ 0.5.

[turn 3] tool_result    {"status":"confirmed","signature":"2mGQ...wisv"}
# Signed and confirmed on devnet. Real signature.

[turn 4] tool_call      {"tool":"get_balance","input":{"address":"Ddfz...c6cN"}}
# Verifies the savings wallet, as the goal instructed.

[turn 4] tool_result    {"lamports":200000000}
# Exactly 0.2 SOL. Goal met.

[turn 5] tool_call      {"tool":"get_balance","input":{"address":"9Sxx...dTf1"}}
# Verifies the operating wallet too.

[turn 5] tool_result    {"lamports":4241965000}
# Dropped by exactly 0.2. Nothing extra moved.

[turn 6] final_report   "The savings wallet now holds 200,000,000 lamports (0.2 SOL)..."
# Honest report. Matches what actually happened on chain.
Enter fullscreen mode Exit fullscreen mode

6 turns. 5 tool calls. 1 confirmed transfer. 0 denials. Goal met exactly.

Annotated run — Trial 2, policy denial

Goal changed to 0.4 SOL. Per-transfer cap tightened to 0.05 SOL.

[turn 3] tool_call      {"tool":"transfer_sol","input":{"lamports":200000000,"to":"Ddfz...c6cN"}}
# Tries to move the full 0.2 SOL in one transfer.

[turn 3] policy_check   {"allowed":false,"reason":"200000000 lamports exceeds the per-transfer cap of 50000000"}
# Refused. 0.2 > 0.05. Nothing signed.

[turn 3] tool_result    {"status":"denied","reason":"..."}
# Model gets the denial as a normal tool result.

[turn 4] tool_call      {"tool":"transfer_sol","input":{"lamports":50000000,"to":"Ddfz...c6cN"}}
# Adjusts. Drops to the cap: 0.05 SOL.

[turn 4] tool_result    {"status":"confirmed","signature":"..."}
# Allowed. Signed.

[turn 5, 6, 7] ...three more 0.05 transfers, all confirmed
# Split the work into pieces.

[turn 12] ...hit MAX_TURNS before writing a final report
Enter fullscreen mode Exit fullscreen mode

12 turns. Hit the limit. Overshot to 0.45 SOL instead of 0.4.

The finding: the per-transfer cap did not stop the agent. It slowed it down. The agent found a path around it — four small transfers instead of one big one. Only the per-run cap bounds total spend.

Lessons learned

1. The policy layer is the heart of the system, not the model.

Everything else is replaceable. The model can change. I went from a local 3B to a 120B on Groq. The prompt can change. The tool-call sequence can change, and it does, run to run. The policy module doesn't care about any of that. It takes a recipient and an amount and returns a decision. Same rules every time. That's the piece you actually trust.

2. MAX_TURNS doesn't care about your goal.

In Trial 2, the agent hit the turn limit mid-plan. It had no way to know it was about to be cut off. It just stopped. No final report, no summary, no "here's where I got to." A turn limit is a hard stop, not a graceful one. In production, you'd want the last turn reserved for a forced report.

3. A per-transfer cap is a rate limit, not a spend limit.

I thought capping transfers at 0.05 SOL would stop the agent from moving 0.2. It didn't. It split the work into four transfers and moved the same amount. The cap changed how the agent spent, not whether it spent. Only the per-run cap bounds total spend. Without it, the per-transfer cap just slows the drain down.

4. Confirmation lag is real.

The model read the savings balance while one transfer was still confirming. Saw less than expected. Sent another. Overshot by 0.05. The model has no built-in sense that a recent write might not have settled yet. Unless you tell it, it treats every read as final.

5. Non-determinism isn't a bug, it's the environment.

I ran the same goal twice in Trial 1. First run: 6 turns, 1 transfer. Second run: 3 turns, 0 transfers — because the goal was already met. Same code, same model, different behavior, both correct. That's not something you fix. It's something you design around. The policy holds regardless of which path the model takes. That's the whole argument.

Open questions

  • Should the policy layer persist spentThisRun across runs, instead of resetting each time?
  • Should transfer_sol wait for finalization rather than confirmation before returning?
  • What happens when the agent needs to move funds to a recipient that isn't the savings wallet? The allowlist is one address today.

These are the edges. They're where the next version starts.


Part of #100DaysOfSolana. Days 92–96 built this stack. Day 97 documented it.


Top comments (1)

Collapse
 
alexshev profile image
Alex Shev •

The deny-by-default policy layer is the part that makes unattended operation discussable rather than merely automated. For transfer-like tools, it is worth logging the evaluated rule set and the pre/post balances with each decision, so an operator can distinguish a model choice from a policy-approved state transition during review.