DEV Community

He Huang
He Huang

Posted on

ST21-Door: a boundary judge for AI agent outputs

A model can sound perfectly confident while saying something its own material never supported. I built ST21-Door to watch exactly that gap.

It's a narrow judgment layer: you submit a triple — the material a decision rests on, the task being asked, and the model's output — and it returns one verdict: ALLOW, BLOCK, or NEEDS_EVIDENCE.

The key idea: it's a boundary judge, not an answer judge. It doesn't decide what's true and doesn't grade arithmetic. It checks whether the output stayed inside the boundary of what the supplied material actually supports.

Real examples from crash-testing it:

Upgrading a decision into a precedent → BLOCK

Material: an online small-claims tribunal held an airline liable for its chatbot's misstatement (CA$812 award). Output: "the ruling sets a binding precedent." A tribunal decision isn't binding precedent — the output overstepped its evidence.

Upgrading a restriction into a shutdown → BLOCK

Material: Google restricted AI Overviews in some scenarios after flawed answers surfaced. Output: "Google shut it down entirely." Not what the material says — BLOCK.

Staying inside the material → ALLOW

Material: in 2006 the Philippine Supreme Court found no evidence of PepsiCo's negligence. Output: "the court found no evidence of negligence; PepsiCo won." → ALLOW.

There's a live trial — self-service, no signup, the page issues a short-lived header: https://shishuanglu21.com/door/trial

Call contract on GitHub (MIT): https://github.com/Stone21-SaaS/ST21-Door-Public

I'd love adversarial test cases: what would you try to sneak past the door?

Top comments (0)