Picture the design review I want us to walk. A designer wants the portfolio to ship tonight. An agent wrote those case studies late last night.
We ask who owns each sentence on the page. That ownership question becomes our whole review gate. The decision owner is the person named on the page.
The consequence is a public claim they may have to defend. The reversible moment is before any write hits the live repo. I will walk this as the review owner.
I am not claiming a real client result. I am showing the protocol I want on the table.
Think of the review like a kitchen pass. The plate should not leave without a ticket. The agent is the cook in this analogy.
You are the pass, not a hungry diner. Smooth fluency is only the plating, not the proof. The source note is the ticket that matters.
Does the source ticket still match the plate? A fluent plate with no ticket should wait.
Freeze the decision
You start this tutorial from an empty folder. You do not open the agent chat yet. You name the decision in a file first.
Can you point to the human decision owner? If you cannot name them, stop the review.
mkdir -p voice-gate/{inbox,record,discard,checks}
printf '%s\n' 'decision: publish case-study sentences' > voice-gate/inbox/decision.txt
printf '%s\n' 'owner: named designer on the page' >> voice-gate/inbox/decision.txt
printf '%s\n' 'live_write: false' >> voice-gate/inbox/decision.txt
grep -q 'live_write: false' voice-gate/inbox/decision.txt && echo 'freeze_ok'
You verify that freeze before you move on. The live write flag must still read false. A true flag means you already lost the reversible moment.
You should see freeze_ok in the terminal. If you do not, fix the file first. You do not negotiate with the unfinished draft.
Split evidence from guesses
I treat the agent paragraphs like unlabeled jars. Some jars hold quotes from the designer notes. Some jars hold smooth guesses with no witness.
Mixing them is how a portfolio starts lying. Would you defend a sentence you did not live?
Write a small record before any rewrite starts. Each sentence needs a source, a speaker, and a status. Use status values of evidence, hypothesis, or discard.
An empty source field means you must stop. A fluent guess dressed up as evidence is worse.
cat > voice-gate/record/sentences.json << 'EOF'
{
"decision_owner": "Amina Rahman",
"page": "portfolio/case-study-checkout",
"sentences": [
{
"id": "s1",
"text": "I interviewed eight cashiers before changing the error copy.",
"source": "research-notes/2026-09-12.md#cashiers",
"speaker": "Amina Rahman",
"status": "evidence"
},
{
"id": "s2",
"text": "The redesign cut checkout time in half.",
"source": "",
"speaker": "agent-draft",
"status": "evidence"
}
]
}
EOF
python3 -c 'import json; json.load(open("voice-gate/record/sentences.json")); print("record_ok")'
Amina is a stand-in name for this protocol. Do not paste a real teammate until they consent. The empty source on s2 is the point.
A bold result with no note is a hypothesis. It must not ship as a research finding. Marking it as evidence is the bug we want the gate to catch.
You should see record_ok before the stop check. A parse error means the record is not ready.
Run the stop checker
Here is the verifier I keep beside the review. It is a proposal script, not a production service. It fails closed when the record is thin.
A missing owner should fail this gate immediately. A hypothesis marked as evidence also fails closed.
cat > voice-gate/checks/stop_gate.py << 'PY'
#!/usr/bin/env python3
"""Proposal checker. Not a live product test."""
import json
import sys
path = sys.argv[1]
data = json.load(open(path, encoding="utf-8"))
errors = []
owner = data.get("decision_owner", "").strip()
if not owner or owner == "agent-draft":
errors.append("missing_owner")
for row in data.get("sentences", []):
sid = row.get("id", "?")
status = row.get("status")
if status not in {"evidence", "hypothesis", "discard"}:
errors.append(f"{sid}: bad_status")
if status == "evidence" and not str(row.get("source", "")).strip():
errors.append(f"{sid}: evidence_without_source")
if row.get("speaker") == "agent-draft" and status == "evidence":
errors.append(f"{sid}: agent_voice_marked_evidence")
if errors:
print("stop")
print("\n".join(errors))
sys.exit(2)
print("review_may_continue")
PY
python3 voice-gate/checks/stop_gate.py voice-gate/record/sentences.json
The command above writes the checker and runs it. On the sample above, you should get stop.
The second sentence still has an empty source. Its speaker is the agent draft, not Amina. That stop is a pass of the protocol.
It is not a failure of the tool. Change s2 to discard or attach a real source.
Run the checker again after you edit it. Only the review_may_continue line lets you draft.
What missing evidence should stop an approval here? An unnamed owner should always stop the approval. A result with no note should stop it too.
A sentence the named person will not say aloud should stop it. What extra information would only add review noise?
A second paraphrase of the same note is noise. A tone score with no rubric is noise. A model ranking you cannot reproduce is noise.
I rehearse three scenarios before I trust the gate. First, a metric sentence with no study note. Its stop condition is an empty source field.
Second comes a quote from a real participant. Stop if that person never consented to be named. Third, an accessibility boast with no pass file.
Stop until checks/a11y.md exists and names a tester. Success here is narrower than a pretty page. Success means the owner can say every kept line.
Success also means every discard still has a reason. A pretty draft without that reason is not success.
Bound the free draft
The outreach brief calls MonkeyCode an open-source project. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
I am not pasting a repo URL I have not opened. Use the official project page you can verify. I use a free model only after the gate passes.
I also use the free server option only then. Both availability claims come from the outreach brief. I am not printing a token quota here.
I am not printing hardware or a duration. Allowances change, so the live product page wins. Copy the live figure into inbox/allowance.txt first.
Do that before you plan any run size. If the page disagrees with me, trust the page.
The server job should not hold publish keys. Give it the notes, not the live repo. Ask for a draft that labels every sentence.
You should refuse any line that invents a metric. A made-up timing cut is a false public claim.
mkdir -p voice-gate/inbox/notes
printf '%s\n' 'Eight cashier interviews informed the error copy. No timing study was logged.' > voice-gate/inbox/notes/cashier-note.md
test -s voice-gate/inbox/notes/cashier-note.md && echo 'notes_ok'
cat > voice-gate/inbox/run-brief.md << 'EOF'
Role: draft assistant. You may not publish.
Inputs: only files in voice-gate/inbox/notes.
Output: voice-gate/record/draft.json
Rules:
- Every sentence needs source, speaker, and status.
- If the note lacks a number, do not invent one.
- Mark guesses as hypothesis.
- If a sentence needs a life you cannot see, use discard.
EOF
test ! -f voice-gate/inbox/live_portfolio_key && echo 'no_live_key_in_workspace'
Create a note the agent is allowed to read. Keep the missing timing study visible in that note. You should see notes_ok before any draft call.
Verify the bound before you paste that brief. No live key should sit in the workspace. The brief forbids invented numbers and fake names.
If setup asks for a deploy token, stop. You are in the wrong stage for that. Remove the token and start the gate over.
I will not claim I timed this run. I have not attached any benchmark to this. If you run it, record your own result.
Put that result beside the allowance file too. Do not borrow my silence as a score.
Keep the discarded lines
People delete the awkward sentence and lose the reason. I keep the discard in a concrete record. That record is the recovery path you will need.
Six months later, someone will ask about the boast. Can you answer them without the discard file?
Set s2 status to discard, then hold the doubt. Hold the doubt even if the line sounds sharp.
python3 - << 'PY'
import json
path = "voice-gate/record/sentences.json"
data = json.load(open(path, encoding="utf-8"))
for row in data["sentences"]:
if row["id"] == "s2" and not str(row.get("source", "")).strip():
row["status"] = "discard"
row["speaker"] = "agent-draft"
json.dump(data, open(path, "w", encoding="utf-8"), indent=2)
held = [r for r in data["sentences"] if r["status"] in {"discard", "hypothesis"}]
json.dump(held, open("voice-gate/discard/held.json", "w", encoding="utf-8"), indent=2)
print("held=%d" % len(held))
PY
python3 voice-gate/checks/stop_gate.py voice-gate/record/sentences.json
test -s voice-gate/discard/held.json && echo 'discard_kept'
Verify that held.json exists after you run this. It should not be empty when a hypothesis existed. An empty discard plus a shiny draft is a smell.
You threw the doubt away far too early. Put the doubt back into the discard record. You should also see review_may_continue after the status edit.
If you still see stop, read the error before you continue. You should see discard_kept beside that continue line.
Review access before ship talk
I do not implement the component in this article. A frontend owner can build the control later. I review the pattern a person must pass.
The publish control needs a visible text name. Color alone must never carry the status meaning. Please use evidence, hypothesis, or discard in text.
Screen reader users need the owner name nearby. It must sit in the same view as approve. If approve sits on another route, say so.
Say that in text, not in a tooltip. Tooltips often fail on touch and on keyboard. I treat that failure as a block, not polish.
Check the reading load before you praise fluency. A case study only a model can parse is not inclusive. Read the kept sentences aloud before any approval.
If you stumble, a participant might stumble too. Cut that sentence and record the cut reason.
python3 - << 'PY'
import json
rows = json.load(open("voice-gate/record/sentences.json", encoding="utf-8"))["sentences"]
long_rows = [r["id"] for r in rows if len(r["text"].split()) > 24]
if long_rows:
print("revise_length " + " ".join(long_rows))
else:
print("a11y_text_pass")
PY
This length check is only a rough proxy. It does not prove contrast or focus order. It also does not prove captions or labels.
Pair it with a short manual accessibility pass. Use keyboard only, then two hundred percent zoom. Add a screen reader spot check on approve.
Write the result in checks/a11y.md before approval. No written file means no approval from me.
Hand the voice back
Approval is not the end of this protocol. The named owner must hear their own sentences. I send the kept lines and the discard file.
I ask which line they would not say in an interview. A no on any line reopens the gate. The agent does not get a vote then.
If the free server run fails, you still have notes. A timeout does not invent a new truth. You do not rebuild truth from the last paragraph.
You return to the notes and rerun the checker. You do not publish what you merely remember.
Who should skip this approach and walk away? Skip it when you need production component code. That build work belongs with the frontend owner.
Skip it if no human will defend the page. Skip it if you need a measured model comparison. This protocol does not produce that kind of comparison.
Skip it for covert collection of private text. Consent lives in the owner field for a reason. No consent means no sentence enters the record.
What I will claim
Would you let a fluent boast survive an interview? Each recommendation here rests on a visible check.
It does not rest on a merely confident vibe. The owner field is the accountability check I trust. The source field splits evidence from a hypothesis.
The stop script acts as the mechanical refusal. The discard file becomes the recovery record later. The accessibility notes stay a manual gate here.
Remember that the word-count script is only a proxy. I am explicit where the evidence is thin. I did not run a live free-server job here.
I did not verify a token total either. Please confirm both on the current MonkeyCode docs. Do that before a study plan depends on them.
If you rehearse this, wait for review_may_continue first. Read the live allowance with your own eyes. Keep the discard file when the draft sounds fluent.
Top comments (0)