DEV Community

Cover image for Agent autonomy gates and action logs in plain Python
Anil Prasad
Anil Prasad

Posted on Originally published at anilsprasad.substack.com

Agent autonomy gates and action logs in plain Python

Gartner predicted on 26 May 2026 that by 2027, 40% of enterprises will demote or decommission autonomous AI agents over governance gaps found only after production incidents. Its recommended fix is to classify agents into four autonomy levels and scale controls to match. This post turns that into code you can run: a manifest, a gate, an append only action log, a circuit breaker, and a reliability check. Standard library only.

The four levels, as data

Level 1 observes (read only). Level 2 advises (drafts, humans act). Level 3 acts with approval (writes after sign off). Level 4 acts autonomously within guardrails. The level belongs in the agent's manifest, where CI can read it.

agent_id: refunds_v3
autonomy_level: 3
allowed_tools: [crm.read, payments.refund]
policy_version: "2027.01"
owner: payments_platform
Enter fullscreen mode Exit fullscreen mode

The gate

Every tool call passes through one function. It decides, and it writes a record whatever the decision.

import hashlib
import json
import time
from dataclasses import dataclass

WRITE_TOOLS = {"payments.refund", "crm.update", "email.send"}

@dataclass
class Agent:
    agent_id: str
    level: int
    allowed_tools: frozenset
    policy_version: str

def input_hash(payload):
    raw = json.dumps(payload, sort_keys=True).encode()
    return hashlib.sha256(raw).hexdigest()

def decide(agent, tool, approver):
    writes = tool in WRITE_TOOLS
    if tool not in agent.allowed_tools:
        return "denied_tool"
    if writes and agent.level < 3:
        return "denied_level"
    if writes and agent.level == 3 and approver is None:
        return "needs_approval"
    return "allowed"

def gate(agent, tool, payload, on_behalf_of, log, approver=None):
    verdict = decide(agent, tool, approver)
    log.append({
        "ts": time.time(),
        "agent_id": agent.agent_id,
        "on_behalf_of": on_behalf_of,
        "autonomy_level": agent.level,
        "tool": tool,
        "input_sha256": input_hash(payload),
        "policy_version": agent.policy_version,
        "verdict": verdict,
        "approver": approver,
    })
    return verdict
Enter fullscreen mode Exit fullscreen mode

Two principals on every record: the agent, and the human or system it acted for. If you record only one, you cannot answer whose authority the action ran under.

The action log

Append only. One line per action, never per session. Output logs tell you what the agent said. This tells you what it was allowed to do, and why.

class ActionLog:
    def __init__(self, path):
        self.path = path

    def append(self, record):
        line = json.dumps(record, sort_keys=True)
        with open(self.path, "a", encoding="utf8") as f:
            f.write(line + "\n")
Enter fullscreen mode Exit fullscreen mode


In production you would ship these lines to storage your agents cannot write to, and chain each record to the hash of the one before it so tampering is detectable. The shape of the record matters more than the store.

The circuit breaker

Gartner's requirements for Level 4 agents include circuit breakers that halt an agent when it crosses a threshold. Here is the smallest useful one: too many non allowed verdicts inside a window, and the agent stops.

class CircuitBreaker:
    def __init__(self, max_denials, window_seconds):
        self.max_denials = max_denials
        self.window = window_seconds
        self.events = []
        self.open = False

    def observe(self, verdict):
        now = time.time()
        self.events = [t for t in self.events if now < t + self.window]
        if verdict != "allowed":
            self.events.append(now)
        if len(self.events) >= self.max_denials:
            self.open = True
        return self.open
Enter fullscreen mode Exit fullscreen mode

When open is true, the agent stops taking actions and pages its owner. It does not retry. An agent that keeps probing for a tool it is not allowed to use is telling you something, and the right response is a human looking at it.

Wiring it together

log = ActionLog("actions.jsonl")
breaker = CircuitBreaker(max_denials=3, window_seconds=300)
agent = Agent("refunds_v3", 3, frozenset({"crm.read", "payments.refund"}), "2027.01")

verdict = gate(agent, "payments.refund", {"order": 4812, "amount": 40},
               on_behalf_of="user:4812", log=log)
if breaker.observe(verdict):
    raise SystemExit("breaker open: agent halted, owner paged")
print(verdict)   # needs_approval, because Level 3 writes require a named approver
Enter fullscreen mode Exit fullscreen mode

The CI check nobody writes

The most common failure is not a bad gate. It is an agent quietly running at Level 4 because nobody declared otherwise. This test fails the build when that happens.

def test_level4_requires_named_approval(manifests):
    for m in manifests:
        if m["autonomy_level"] == 4:
            assert m.get("level4_approved_by"), m["agent_id"]

def test_write_tools_match_level(manifests):
    for m in manifests:
        writes = set(m["allowed_tools"]) & WRITE_TOOLS
        if writes:
            assert m["autonomy_level"] >= 3, m["agent_id"]
Enter fullscreen mode Exit fullscreen mode

Measuring reliability instead of accuracy

A study on arXiv (2603.29231, March 2026, 23,392 episodes) found that capability rankings and reliability rankings of agents diverged. Your dashboard accuracy is what the agent did once. Promotion to a higher level should depend on what it does every time.

def passes_every_time(task, run, variants, check, n=8):
    return all(check(task, run(v)) for v in variants(task, n))

def reliability(tasks, run, variants, check):
    passed = sum(1 for t in tasks if passes_every_time(t, run, variants, check))
    return passed / len(tasks)
Enter fullscreen mode Exit fullscreen mode

variants should produce honest variations: rephrased requests, reordered records, realistic missing fields. Eight identical calls measures your cache. Eight honest variations measures your system.


A promotion rule you can defend

Put these together and you get a rule an auditor can follow. An agent moves up one level when its reliability over the last N tasks clears a threshold you set, its breaker has not opened in that period, and a named owner signs the manifest change. It moves down automatically when the breaker opens. Every step is backed by records the agent could not edit.

What this does not cover

The numbers above, three denials in five minutes and eight variations per task, are starting points I would expect to tune. I do not know of published evidence for specific values, and I would distrust anyone who quotes one without their own data behind it.

Identity federation across organizations, which is where the A2A protocol and your identity provider meet. Prompt injection defence inside tool outputs. The OWASP Top 10 for Agentic Applications (December 2025) is the right checklist for both. The gate above is the part most teams skip, and it is the part every incident review asks for first.

If you run agents in production: what threshold do you use before letting one write without approval, and was it chosen from data or from a meeting?

humanwritten #expertisefromfield

Anil Prasad has spent 25 years building and running production data and machine learning systems. He is the founder of Ambharii Labs, where he designs agent infrastructure for healthcare revenue cycle work. That includes a runtime that scores every model output before release and an append only log of every agent decision. He writes about the unglamorous parts of AI engineering that decide whether a system survives its first incident. The code in this post is illustrative and uses only the Python standard library. Questions and pull apart critiques are welcome in the comments.

Top comments (0)