The rain in Halifax had already fogged the library windows, and my laptop was sitting at four percent. I had a two-sentence source I had written for myself, a model draft that sounded friendly, and a lab report that wanted to leave with the bus. The draft felt like my voice. Then I searched for "local slope" and found a full sentence living in the prose, unmarked, as if I had thought it.
Had I written that line, or had the draft borrowed it while I watched the battery icon? That is the bug I sat down to catch. Not a leaderboard story, and not a tour of product screens. A note that steals a sentence from the source you pasted above the prompt, then smiles.
The mess I walked in with
I keep a small journal for ML foundations so later-me can see what I actually understood, not what a model performed on my behalf. The failure has a plain name in my notes: an attribution leak. You ask for a paraphrase, most of the words move, and a six-word run stays put. A week later the note reads as yours.
Would I put that note in a portfolio repo? No. A page can look finished and still hide a sentence I do not own. People are busy showing off profile pages this week, and I get the itch, but a pretty page would not have caught this. A boring checker might.
I am not building a plagiarism court. I am also not interested in helping anyone wash copied coursework until a detector shrugs. If the source is an assignment, a paper you do not have the right to remix, or someone else's unpublished notes, stop. This lab uses a fixture I wrote, so the only thing under test is the fence.
The one question
Can a standard-library script show me the stolen sentence before I forget it was stolen? That is the whole goal. Prerequisites stay dull on purpose. You want Python 3.11 or newer, a terminal, and no third-party packages.
I checked the version assumption because the type hints in the script use list[str]. If python3 --version prints something older, do not "fix" the lab with a surprise install. Use a newer interpreter, or read the file as a sketch you will retype. I also wanted the run to survive a dying laptop. A free shell with Python is enough, and a GPU never enters the story.
Where a hosted seat actually fits
Here is the only product part, and I want it beside the limit rather than in the title. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The availability claims I will repeat are the ones I was given: free model access, and a free server option. I will not name models, quotas, regions, hardware, or a duration.
I did not measure those, and a student write-up that invents them is just another unmarked sentence. Free model access, in this workflow, is only the drafting seat. I can ask for a paraphrase of my own fixture, and I can ask that any leftover verbatim sit on lines starting with >. Then I leave. The judgment happens in the script, because a hosted draft is a guest and the fence is the door.
The free server option is the other seat. If the library laptop is about to sleep, a small remote shell that already has Python can run the same file. If a spare shell is what you are missing, that is the option I would try, after a local copy exists. I would not store the only copy of the lab on a machine I do not administer. A free box can disappear, and I have no promise that it will not.
If you need a private exam or another person's notes kept off third-party prompts, skip that drafting seat. Type the fixture yourself. The checker still teaches the same lesson.
The fence
The rule is almost rude. Normalize to lowercase word tokens, then build every run of six tokens from the source and from the unmarked prose. If those sets intersect, the prose stole a sentence-shaped piece, even when you wrapped the borrow in new words. Lines that start with > are not prose leaks. They are claims of quotation, so each one must be a normalized substring of the source.
A fake quote fails even when the prose is clean. An empty draft fails closed. I would rather a noisy red light than a green light on a blank file. Why six tokens, and not a threshold I copied from a paper? Because this is a teaching knob. Shorter windows start yelling about ordinary English like "of a loss". Longer windows let a five-word theft stroll past. My stolen sentence was longer than six, so the fixture fails for a reason I can point at.
#!/usr/bin/env python3
"""Fence unmarked verbatim runs before they enter a lab note."""
import re
import sys
from pathlib import Path
WINDOW = 6
def tokens(text: str) -> list[str]:
return re.findall(r"[a-z0-9]+", text.lower())
def windows(toks: list[str], n: int) -> set[tuple[str, ...]]:
if len(toks) < n:
return set()
return {tuple(toks[i : i + n]) for i in range(len(toks) - n + 1)}
def main() -> int:
if len(sys.argv) != 3:
print("usage: python3 quote_fence.py SOURCE DRAFT")
return 2
src_path, draft_path = Path(sys.argv[1]), Path(sys.argv[2])
if not src_path.is_file() or not draft_path.is_file():
print("status=fail")
print("reason=missing_file")
return 1
source = src_path.read_text(encoding="utf-8")
draft = draft_path.read_text(encoding="utf-8")
if not draft.strip():
print("status=fail")
print("reason=empty_draft")
return 1
source_norm = " ".join(tokens(source))
src_wins = windows(tokens(source), WINDOW)
prose_lines, quote_lines = [], []
for line in draft.splitlines():
if line.startswith(">"):
quote_lines.append(line[1:].strip())
else:
prose_lines.append(line)
leaks = windows(tokens("\n".join(prose_lines)), WINDOW) & src_wins
bad_quotes = []
for line in quote_lines:
norm = " ".join(tokens(line))
if not norm or norm not in source_norm:
bad_quotes.append(norm or "<empty>")
print(f"source_windows={len(src_wins)}")
print(f"prose_leaks={len(leaks)}")
if leaks:
print("leak_example=" + " ".join(sorted(leaks)[0]))
print(f"bad_quotes={len(bad_quotes)}")
if bad_quotes:
print("bad_quote_example=" + sorted(bad_quotes)[0])
failed = bool(leaks or bad_quotes)
print("status=fail" if failed else "status=pass")
return 1 if failed else 0
if __name__ == "__main__":
raise SystemExit(main())
Save that as quote_fence.py. The source fixture is mine, two sentences about a gradient step, so this write-up does not borrow a textbook page. The leak draft repeats one sentence with no > mark. The fake-quote draft speaks in my own words, then cites a sentence the source never said. The good draft quotes the real sentence and paraphrases the rest without a six-token echo. The empty file is the input that should not look like success.
mkdir -p quote-fence && cd quote-fence
python3 --version
cat > source.txt << 'EOF'
A gradient step follows the local slope of a loss surface. If the step is too large, the next point can jump past the valley and the loss can rise.
EOF
cat > draft_leak.txt << 'EOF'
I think optimization is just walking downhill.
A gradient step follows the local slope of a loss surface.
EOF
cat > draft_fake_quote.txt << 'EOF'
I picture a loss surface as a valley in fog.
> If the step is too huge, magic happens and loss always falls.
EOF
cat > draft_ok.txt << 'EOF'
I picture a loss surface as a valley in fog.
> A gradient step follows the local slope of a loss surface.
A large step can overshoot, so the next loss is not guaranteed to drop.
EOF
: > draft_empty.txt
python3 quote_fence.py source.txt draft_leak.txt ; echo "exit=$?"
python3 quote_fence.py source.txt draft_fake_quote.txt ; echo "exit=$?"
python3 quote_fence.py source.txt draft_ok.txt ; echo "exit=$?"
python3 quote_fence.py source.txt draft_empty.txt ; echo "exit=$?"
python3 quote_fence.py source.txt missing.txt ; echo "exit=$?"
I derived the expected lines by applying those token rules on paper. Treat them as the spec to confirm when you run the file, not as a log I copied from a server. If your interpreter surprises you, the fixture wins the argument, not my memory.
The leak draft should print source_windows=25, then prose_leaks=6, then leak_example=a gradient step follows the local, then bad_quotes=0, then status=fail. The exit code should be 1. The fake quote should print prose_leaks=0 and bad_quotes=1, with bad_quote_example=if the step is too huge magic happens and loss always falls, then status=fail.
The good draft should print prose_leaks=0, bad_quotes=0, and status=pass, with exit 0. The empty draft should print status=fail and reason=empty_draft. A missing path should print status=fail and reason=missing_file, and it must not pretend the note was clean.
What the run taught me
The interesting failure is not the script crashing. It is the green light you can fake by accident. Mark every line with > and the prose leak count drops to zero, even though you hid the theft inside the fence. The substring check still demands that the hidden line exist in the source, so a pure invention stays red.
A real sentence, wrongly marked as a casual quote, goes green. The tool believes your mark. Should it? Only if you treat > as a promise you would defend in office hours, not as camouflage. The mark is not a citation style. It is a flag the script can see.
The miss in the other direction matters just as much. Change "large" to "big", shuffle the clauses, and a six-token window may pass a note that is still a close rewrite. That is not the script being kind. It is the script being blind to meaning. If your course forbids close paraphrase, a pass here is not permission. I want that limit in the same breath as the happy path, because a checker that only celebrates will train you to game it.
A few slips showed up while I sketched the fixtures. Forgetting .lower() makes Gradient and gradient look like different words, so a leak survives. Dropping the window to three turns harmless phrases into sirens. Reading status=pass as "the model told the truth" confuses a string match with an understanding check. Copying the only fixture onto a free server, then closing the laptop, confuses convenience with a backup.
After the battery warning
What should you understand once those commands settle? A draft can sound like you and still carry a sentence-shaped borrow. A quote mark is a claim, and claims can be checked against the source you say you used. A pass is a narrow event: no six-token prose overlap, and every marked quote is actually in the source.
It is not originality. It is not correctness. It is not a citation style, and it is not a grade.
I would not use this approach if you need semantic plagiarism detection, legal clearance, or an academic-integrity verdict. I would not point it at other people's private writing, and I would not feed a hosted model material you are not allowed to share. I would not treat free model access as a library, or a free server as an archive. Those seats are optional convenience around a script that should still run on your machine if the offer changes. That portability is the part I trust.
Which file did you expect to fail, and which failure did you mis-rank? I had the fake quote feeling more wrong than the verbatim leak, and the sets disagreed with my gut until I wrote them out. If you can shrink the good draft until it fails for a reason the window cannot justify, that counterexample is the extension I want. Bring the fixture, not the applause.
Top comments (0)