A free model will happily plan a bigger change than the one you asked for. I do not spend that call until a local file says what "enough" means.
You know the feeling? You typed "rename the helper," and the plan came back with a new HTTP client, a lockfile edit, and a path you never named. Who approved that growth? Not me. Not the ticket. The model did, because nothing on disk was allowed to say no.
This is a from-zero tutorial for that no. Intent file, plan fixture, checker, then a verification command after every stage. The samples are fixtures you can read in a terminal. I am not describing a production outage, a customer, or a measured quota.
What I am actually refusing
I want one rule you can reuse tomorrow. The human writes the bound. The model may propose work inside the bound. If the plan adds a path, a host, a dependency, or a command, the job fails closed before any free server sees it.
Is that a sandbox? No. Is it code review? Also no. It is a small refusal you can run before the interesting part starts.
A DEV thread this week poked at a sharper fear: a model that notices a task is aimed at someone else's system and stays quiet. I am not re-running that experiment, and I am not treating a title as a measurement. I am asking a narrower question. What does your runner do when the plan grows past the sentence you wrote?
Step 1. Write the intent before you open a chat
Start empty. I do not want a chat transcript to be the spec.
mkdir -p intent-gate/fixtures
cd intent-gate
Put this in intent.json.
{
"intent_id": "rename-helper-014",
"summary": "Rename format_name to format_label in pkg/names.py only.",
"allowed_paths": ["pkg/names.py"],
"allowed_hosts": [],
"allowed_deps": [],
"allowed_commands": ["python -m pytest tests/test_names.py"],
"forbidden": ["network", "lockfile", "shell_interpolation"]
}
Verify the file is actually JSON.
python -c "import json; json.load(open('intent.json')); print('intent_ok')"
You want intent_ok. Broken JSON? Stop. Why would I let a model parse the bound I failed to parse myself?
Step 2. Pin the checker before you pin a model
Save widen_check.py. It loads both files. It exits 0 only when every path, host, dependency, and command in the plan is already listed, and the intent id matches.
#!/usr/bin/env python3
"""Local plan gate. Unexecuted against a live server in this draft."""
import json
import sys
def load(path):
with open(path, encoding="utf-8") as handle:
return json.load(handle)
def main():
intent = load(sys.argv[1])
plan = load(sys.argv[2])
reasons = []
pairs = (
("paths", "allowed_paths"),
("hosts", "allowed_hosts"),
("deps", "allowed_deps"),
("commands", "allowed_commands"),
)
for key, field in pairs:
allowed = set(intent.get(field, []))
for item in plan.get(key, []):
if item not in allowed:
reasons.append(f"widen:{key}:{item}")
if plan.get("intent_id") != intent.get("intent_id"):
reasons.append("intent_id_mismatch")
if reasons:
print("reject")
for reason in reasons:
print(reason)
return 2
print("accept")
print(intent["intent_id"])
return 0
if __name__ == "__main__":
sys.exit(main())
Compile it. Still no network.
python -m py_compile widen_check.py && echo compile_ok
compile_ok means the gate parses. It does not mean the logic is right. That is the next two steps. Skip them and you only proved Python can read a file.
Step 3. Make the wide plan fail in public
Write fixtures/plan_wide.json.
{
"intent_id": "rename-helper-014",
"paths": ["pkg/names.py", "pkg/net.py"],
"hosts": ["example.invalid"],
"deps": ["requests"],
"commands": [
"python -m pytest tests/test_names.py",
"curl https://example.invalid"
]
}
Run it.
python widen_check.py intent.json fixtures/plan_wide.json; echo exit:$?
I expect reject, then widen:paths:pkg/net.py, widen:hosts:example.invalid, widen:deps:requests, and a widened command line, then exit:2.
Did one of those lines go missing? Then the checker is softer than the fixture, and the fixture is the spec. Fix the script. Do not "try it on the model and see." Seeing is how extra hosts survive.
Step 4. Make the narrow plan pass, and keep it narrow
Write fixtures/plan_fit.json.
{
"intent_id": "rename-helper-014",
"paths": ["pkg/names.py"],
"hosts": [],
"deps": [],
"commands": ["python -m pytest tests/test_names.py"]
}
python widen_check.py intent.json fixtures/plan_fit.json; echo exit:$?
Expected text is accept, then rename-helper-014, then exit:0.
Tempted to delete the wide fixture once this looks green? Don't. A gate that only meets friendly plans is a diary entry. I keep both files, and I rerun both commands whenever the checker changes.
Step 5. Only now do a free model or a free server enter
Disclosure: This article was prepared as part of MonkeyCode's product outreach. This is the first point where that product sits in the method. The claim I will repeat is small: free model access, and a free server option. I am not naming a model. I am not stating a token count, a machine size, or how long either offer lasts. Those numbers go stale. Read the project docs on the day you run the gate, and trust the docs over any post, including this one.
What do I send? The intent id, the summary, and the allowlists. Not the tree. Not environment files. Not a shell line built by gluing tool arguments together.
The reply has to become fixtures/plan_from_model.json before it can matter. Prose is not a plan. If the model chats, I fail closed. I do not scrape a command out of a paragraph.
python widen_check.py intent.json fixtures/plan_from_model.json; echo exit:$?
Exit 2 means I do not "repair it on the server." I either edit intent.json because I truly want the wider job, or I drop the plan. Exit 0 means the free server may run the one command already listed in allowed_commands.
Would I let that server invent a second command? No. The command was frozen in step 1. A free minute is not a new spec.
Read the result like a table, not a vibe
| Plan signal | Local result | What I do next |
|---|---|---|
| Extra path, host, dependency, or command | exit 2 | No server minute |
| Intent id does not match | exit 2 | Discard; do not retry the same file |
| Exact allowlist match | exit 0 | Optional run of the listed command |
| Prose instead of JSON | no score | Fail closed; do not parse English for a shell line |
Where this breaks, and who should walk away
String equality is blunt. pkg/names.py and ./pkg/names.py are different strings here, and a relative path can still sneak past a tired reviewer. Normalize before you share this gate. I left normalization out so the fixture stays readable. That is a limitation, not a feature I want you to miss.
Inside an allowed file, the checker is blind. A wrong rename still looks like accept. An already-allowed dependency can still be compromised. Nothing here proves a free server is empty, isolated, or logged. If the job touches production secrets, payments, or another person's systems, do not use a tutorial script as the control. You want a real sandbox, a reviewer, and an owner who can refuse.
Skip it if you wanted an unbounded generation pass. The method starts with a file you were willing to type by hand. No file, no free call. Why spend a hosted minute to discover you never decided the bound?
After the exit codes match
I would commit the checker and both fixtures, and I would make the wide plan keep failing in CI. The fit plan must keep passing. That pair is the artifact. The model is a guest.
If you want to try the same order, start from the reject fixture, not from a signup page. When both exit codes match this tutorial, then check MonkeyCode's current free model access and free server option and decide whether they fit the job you already wrote down.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.