Every writer who pastes drafts into an AI humanizer eventually asks the same thing: how do I know the text actually sounds human before I send it? I built a small detector to answer that, and the interesting part was not the code — it was watching where it confidently fooled itself.
Here is the whole thing, runnable as-is.
import re
from collections import Counter
def sentences(text):
# split on . ! ? followed by whitespace; keeps the delimiter
return [s.strip() for s in re.split(r'(?<=[.!?])\s+', text) if s.strip()]
def burstiness(text):
lens = [len(s.split()) for s in sentences(text)]
if len(lens) < 3:
return None # too short to mean anything
mean = sum(lens) / len(lens)
var = sum((x - mean) ** 2 for x in lens) / len(lens)
return var ** 0.5 / mean # coefficient of variation
def repeated_ngrams(text, n=3, threshold=2):
words = re.findall(r"[a-z']+", text.lower())
grams = [" ".join(words[i:i+n]) for i in range(len(words)-n+1)]
return {g: c for g, c in Counter(grams).items() if c >= threshold}
def score(text):
b = burstiness(text)
rep = repeated_ngrams(text)
# low burstiness + repeated phrases => suspicious
if b is None:
return None
return round((1 - min(b, 1)) * 0.6 + min(len(rep), 10) / 10 * 0.4, 3)
if __name__ == "__main__":
human = "I shipped the proxy on a Tuesday. Two days later the stream broke. Turned out the client buffered wrong. Fixed it by Friday."
robot = "The implementation of the proxy was completed successfully. The system was designed to be robust. The architecture ensures reliability. The solution provides scalability. The approach demonstrates effectiveness."
print("human:", score(human)) # usually higher (varied lengths)
print("robot:", score(robot)) # usually lower (uniform, repetitive)
Run it and you will see the robot sample scores lower almost every time. Good enough to sort a pile of drafts, right? Three things broke that assumption in production.
Pitfall 1: short texts are noise. burstiness returns None under 3 sentences, and even at 3–4 sentences the coefficient of variation swings wildly. A two-sentence human reply can look more "AI" than a ten-sentence marketing page. I had to treat anything under ~40 words as "unknown" instead of scoring it.
Pitfall 2: genre beats author. Legal summaries, release notes, and API docs are supposed to be uniform. My detector flagged real human technical writing as AI because the domain is naturally low-burstiness. Burstiness alone is a genre signal, not an authorship signal.
Pitfall 3: the detector is the easiest thing to game. Swapping two sentences, splitting one long one into two, or changing "utilize" to "use" drops the repeated-ngram count to zero. That is exactly what a cheap humanizer does — it moves the score without moving the prose.
So detection helped me triage, but it never told me the text was good. The part that actually changed how the output reads was a fixed set of rewrite moves: vary sentence openings, cut the second "that", replace noun-stacks with verbs, and break the rhythm every third sentence. We baked those into a tool we use daily — ShipCopy — but the patterns work fine by hand once you have seen them fail a detector a few times.
The takeaway: a 40-line detector is a fine smoke test. It is not a judge of quality, and anyone who promises the text is human is selling the score, not the sentence.
Top comments (0)