DEV Community

Cover image for The Agent Shouldn't Be Able to Approve Its Own Rules
Vlad Zoff
Vlad Zoff

Posted on

The Agent Shouldn't Be Able to Approve Its Own Rules

Once a coding agent can change the architecture, changing the rules that protect it becomes a different kind of operation.

One of the less obvious problems I've run into with coding agents isn't that they make bad changes.

It's that sometimes they make a perfectly reasonable change that invalidates one of the rules I'm using to check the project.

That's a different problem.

Imagine a project has this rule:

All payment-provider access must go through PaymentAdapter.
Enter fullscreen mode Exit fullscreen mode

An agent is asked to add support for another payment provider.

It looks at the existing code and decides that the current adapter isn't quite right. It wants to introduce a new abstraction.

The resulting change crosses a boundary that the project currently protects.

The checker reports a violation.

So what should happen next?

The obvious automation is:

  1. agent changes the code;
  2. check fails;
  3. agent changes the rule;
  4. check passes.

Technically, everything worked.

But the system has a problem: The thing being checked was allowed to change the conditions of the check.

Diagram separating implementation changes from policy changes

That made me much more interested in the difference between changing the code and changing the rules around the code.

The dangerous loop

The simplest version looks like this:

agent
  ↓
change code
  ↓
verification
  ↓
failure
  ↓
agent changes policy
  ↓
verification
  ↓
pass
Enter fullscreen mode Exit fullscreen mode

There is nothing obviously broken here.

The agent might even have a good reason for changing the policy.

The problem is that the verification boundary has disappeared.

A policy is supposed to tell the system what has to remain true.

Agent change and policy proposal passing through separate verification and approval paths

If the same actor that made the change can also redefine what "true" means, a successful verification doesn't tell you very much.

It's just a moving target.

And this isn't specific to architecture.

You could do the same thing with:

maximum service size = 300 lines
Enter fullscreen mode Exit fullscreen mode

The agent produces a 420-line service.

The check fails.

The agent changes the limit to 500.

The check passes.

Or:

module A cannot import module B
Enter fullscreen mode Exit fullscreen mode

The agent needs the dependency.

It changes the rule.

The import is now allowed.

Again, maybe that's the right architectural decision.

But those are two separate decisions:

  • The implementation changed.
  • The policy changed.

They shouldn't become one operation just because the same agent proposed both.

Not every policy violation is actually a mistake

This is where it gets more interesting.

I don't want an architecture checker that treats every violation as proof that the agent did something wrong.

Sometimes the agent really should change the architecture.

Suppose an application has:

Checkout
   ↓
PaymentAdapter
   ↓
Stripe
Enter fullscreen mode Exit fullscreen mode

And I decide that the product is getting large enough that payment workflows deserve their own domain boundary:

Checkout
   ↓
PaymentService
   ↓
PaymentAdapter
   ↓
Stripe
Enter fullscreen mode Exit fullscreen mode

The change introduces new files.

Some imports move.

Some old boundaries disappear.

New ones appear.

A strict checker could report a pile of violations.

That doesn't mean the change is bad.

It means the current policy describes the old architecture.

This distinction matters.

A policy isn't supposed to prevent architecture from ever changing.

It is supposed to make architecture changes explicit.

Proposal and approval are different things

This led me to a fairly simple rule: An agent should be able to propose a policy change, but it shouldn't automatically be able to approve that policy change.

For example:

Current policy:

Payment provider access must go through PaymentAdapter.
Enter fullscreen mode Exit fullscreen mode

The agent can say:

Proposed change:

Allow PaymentService to depend directly on a new internal
PaymentProvider interface.

Reason:
The current adapter boundary prevents the new workflow
from sharing transaction state correctly.
Enter fullscreen mode Exit fullscreen mode

That's useful.

The agent has done the hard reasoning.

It has identified the existing constraint.

It has explained why the constraint may no longer fit.

But the proposal should remain a proposal.

The important part is that the authority approving the policy change is separate from the agent that authored the change.

Otherwise the system can silently move the goalposts.

This is the same reason I don't want approval in the prompt

A prompt can say:

Never deploy without approval.
Enter fullscreen mode Exit fullscreen mode

That's useful instruction.

It's not much of a control if the same process can modify the configuration that defines what counts as an approved deployment.

The more autonomous the agent becomes, the more these distinctions move out of the prompt and into the environment around it.

The agent should be able to reason about the policy.

It should be able to request a policy change.

It should be able to explain the change.

The actual enforcement shouldn't depend on the agent remembering to follow its own instructions.

A policy change should leave a trail

Once policy becomes a real project artifact, another problem appears.

You need to know what the policy was when the original change was checked.

Otherwise you can end up with a strange situation where today's successful verification only makes sense because yesterday's policy was replaced.

That's why I like keeping policy changes explicit and versioned.

Something like:

project state
    ↓
policy revision 17
    ↓
agent proposes architecture change
    ↓
verification fails under policy revision 17
    ↓
policy proposal
    ↓
approval
    ↓
policy revision 18
    ↓
verify change again
Enter fullscreen mode Exit fullscreen mode

Now there are two separate facts:

  1. The change passed under policy revision 18.
  2. Policy revision 18 was itself approved.

That's much more useful than simply seeing a green check.

The old policy still matters

There's another subtle point here.

Suppose an agent changes the rule and then verifies the same diff against the new rule.

You can no longer tell whether the original change violated the previous policy.

So I want the system to preserve the distinction between:

what the project allowed before the change
Enter fullscreen mode Exit fullscreen mode

and:

what the project allows after the change
Enter fullscreen mode Exit fullscreen mode

This becomes especially useful when investigating a change later.

You can ask:

  • Why did this dependency become allowed?
  • Was it always allowed?
  • Was there an explicit exception?
  • Did a policy revision happen at the same time?
  • Who approved it?
  • What evidence led to the change?

Those questions are much harder to answer when policy is just another mutable config file.

Temporary exceptions are different again

Sometimes the policy is fine.

The violation is temporary.

For example:

All persistence access must go through Repository.
Enter fullscreen mode Exit fullscreen mode

But I'm in the middle of a migration.

I don't want to remove the rule.

I just need one known exception for two weeks.

That's not really a policy change.

It's a waiver.

And I think treating it as a different object makes the whole system easier to reason about.

A useful waiver has at least:

owner
reason
scope
expiry
Enter fullscreen mode Exit fullscreen mode

So instead of:

remove the rule
Enter fullscreen mode Exit fullscreen mode

you get:

waive this finding until 2026-12-28

owner: platform-team
reason: repository migration
Enter fullscreen mode Exit fullscreen mode

The rule stays.

The exception expires.

That's a very different thing from changing the architecture policy permanently.

Baselines solve a different problem

I also don't want to confuse waivers with baselines.

A baseline answers:

This violation already existed.
Enter fullscreen mode Exit fullscreen mode

A waiver answers:

This active violation is intentionally allowed for a limited time.
Enter fullscreen mode Exit fullscreen mode

And a policy change answers:

We changed what the project considers acceptable.
Enter fullscreen mode Exit fullscreen mode

Those are three different states.

That separation might sound overly precise.

Three distinct concepts: baseline, temporary waiver and policy change

In practice, it makes the tool much more useful.

A real codebase can have old architectural debt.

It can have temporary migration exceptions.

And it can deliberately evolve its architecture.

If all three become "ignore this finding", you lose important information.

What the agent should actually do

This doesn't mean agents have to stop making architectural changes.

Quite the opposite.

I want them to do more.

A useful agent flow could look like this:

requirement
    ↓
agent plans change
    ↓
implementation
    ↓
deterministic verification
    ↓
failure
    ↓
agent explains why
    ↓
policy proposal / waiver proposal
    ↓
approval
    ↓
verification against new state
Enter fullscreen mode Exit fullscreen mode

The agent can drive most of that workflow.

It can inspect the repository.

It can understand the requirement.

It can implement the refactor.

It can identify the policy that stopped the change.

It can prepare the evidence for a policy proposal.

It can even tell me that the current architecture appears to be the problem.

What it shouldn't get is an invisible path from:

my change failed
Enter fullscreen mode Exit fullscreen mode

to:

therefore my own change is now allowed
Enter fullscreen mode Exit fullscreen mode

This also changes what "autonomous" means

I've started thinking that autonomy isn't really one switch.

An agent can have permission to:

  • read the repository;
  • modify source files;
  • run tests;
  • inspect dependencies;
  • propose policy changes;

without having permission to:

  • approve those policy changes;
  • remove its own enforcement;
  • extend a temporary waiver indefinitely;
  • change the authority that verifies it.

That gives you a more useful permission model than simply:

autonomous = yes/no
Enter fullscreen mode Exit fullscreen mode

Different operations can have different authorities.

And that matters more once the agent is running for a long time or can delegate work to other agents.

The verifier should not care who made the code change

There's another distinction here that I find useful.

Verification should primarily answer:

Does the current change fit the current approved policy?
Enter fullscreen mode Exit fullscreen mode

It doesn't need to decide whether the author was:

human
Claude
Codex
Cursor
another agent
automation
Enter fullscreen mode Exit fullscreen mode

That's a separate concern.

The verifier checks the state.

The policy layer defines the constraint.

The authority layer controls who can change the constraint.

Keeping those pieces separate makes the system much easier to reason about.

This is what I started building into Guard

This distinction ended up affecting the design of Codapult Guard quite a bit.

Guard already had the idea of:

facts
  ↓
policy
  ↓
verification
Enter fullscreen mode Exit fullscreen mode

But that isn't enough when policy itself can change.

So I started treating policy changes as first-class operations.

Guard can discover project facts and prepare proposals.

A project can explicitly approve them.

For protected projects, policy approval can require a distinct actor rather than allowing the authoring agent to approve its own proposal.

The same idea now applies to waivers.

A waiver is not just "ignore this finding forever."

It has an owner, reason and expiry date.

When the expiry is reached, the finding becomes active again.

That gives the project three useful things:

policy
exception
history
Enter fullscreen mode Exit fullscreen mode

instead of one growing collection of ignored warnings.

It also makes agent integration more useful

MCP makes this distinction especially important.

An agent can ask Guard:

What policy applies here?
Enter fullscreen mode Exit fullscreen mode

or:

What did this change affect?
Enter fullscreen mode Exit fullscreen mode

or:

Why did verification fail?
Enter fullscreen mode Exit fullscreen mode

or:

What policy change would be needed for this refactor?
Enter fullscreen mode Exit fullscreen mode

Those are useful questions.

But I don't want the agent to receive:

Change policy until verification passes.
Enter fullscreen mode Exit fullscreen mode

That isn't really verification anymore.

The MCP layer should expose the information and operations the agent needs without quietly giving it the authority to redefine the system around itself.

That's a much more interesting role for project-aware tooling than simply adding another set of commands to an agent.

The rule can change. That isn't the problem.

I don't think architectural rules should be permanent.

Projects change.

Requirements change.

Teams change.

The shape of the system changes.

A rule that made perfect sense six months ago might become actively harmful.

The mistake is treating policy change as an implementation detail.

Changing:

src/payments/StripeAdapter.ts
Enter fullscreen mode Exit fullscreen mode

and changing:

payment-provider-access
Enter fullscreen mode Exit fullscreen mode

are fundamentally different operations.

One changes the system.

The other changes what the system is allowed to become.

Once an agent can perform both, they need separate boundaries.

I'm less interested in stopping agents than in making their authority explicit

The goal isn't to make agents weaker.

I actually want agents to make bigger changes.

But bigger changes require clearer boundaries.

The more of the implementation I delegate, the more I care about questions like:

  • What can the agent change?
  • What can it propose?
  • What can it approve?
  • What evidence does it have to provide?
  • What policy was active when the change was checked?
  • Was the exception temporary?
  • Who approved the exception?
  • Can the agent change the thing that verifies it?

Those questions aren't really about whether the model is smart enough.

They're about the architecture around the model.

And that's probably the part that becomes more important as coding agents become capable of doing more work without waiting for a human after every step.

I don't want the agent to be afraid of the rules.

I want it to understand them.

I want it to be able to challenge them.

I want it to be able to propose better ones.

But I don't want it to be the final authority on whether its own change should become the new rule.

The code can change.

The architecture can change.

Even the policy can change.

Those changes just shouldn't all happen under the same authority.

Top comments (2)

Collapse
 
anp2network profile image
ANP2 Network •

The proposer/approver separation needs an operational check: count the identities actually exercising each role. In a public signed-event ledger I scanned, 1,482 verdicts were signed by just two keys. The verifier and artifact-producer key sets overlapped by one key, which graded its own delivery. Separate role fields do not establish separate actors.

A cheap audit would count distinct approver IDs in approval records, then count how many also authored the change they approved. Designing separation and checking whether it still holds are separate jobs.

The revision example needs a binding between the revision identifier and its contents. Across 31 keys publishing capability declarations, five changed their declaration bodies. Two kept version at "1.0" through those changes. One removed its pricing description 27 seconds later without changing the version. A reader using that label alone for cache invalidation would miss the edit indefinitely. Record the policy's content hash in each check result, so "passed under revision 18" identifies the exact policy evaluated.

Which side holds that reference matters. An immutable check result pointing to a policy hash preserves what governed that check. A policy repository maintaining the reverse mapping, "these checks used this policy", leaves that mapping exposed to whoever can rewrite the repository. The ledger shows a related limit on traces: children reference parents, while timestamps are self-reported with no independent clock. Eleven work-time entries preceded their own job acceptance. That leaves room for a backdated acceptance to displace an earlier, honestly dated one. A signed reference alone cannot establish chronology.

Reachability failed too. Of 1,419 tasks carrying verdicts, 1,400 of them, or 98.6%, had an empty public consensus field. The evidence was stored, and the published view gave no path to it. A green check needs a traversable link to the approval it rests on.

Does your mechanism retain enough identity information to count, afterward, whether approvals actually came from actors distinct from the authors of the changes they approved?

Collapse
 
hellomathieup profile image
Mathieu Poli •

Seen a smaller version of this: the agent "fixing" a failing test by editing the test. Same thing, the check stops meaning anything. Do you keep the rules somewhere the agent can read but not write?