AI detectors are usually tested with a simple question: is this text human-written or AI-generated?
But there’s another question that deserves attention: does the genre of human writing affect the result?
A personal email, a news article, a technical document, and a short story can all be completely human-written while having very different sentence structures, vocabulary, tone, and predictability.
That makes a cross-genre AI detector test particularly interesting.
The Experiment: Test Only Verified Human Writing
For this experiment, I’d start with known-provenance human samples. That means every document should have a reliable history showing that it was written by a person rather than generated or rewritten using generative AI.
I’d divide the dataset into five categories:
Academic essays — Formal writing with structured arguments, citations, and relatively predictable organization.
Professional emails — Shorter, direct communication that may use common phrases and standard formatting.
Journalism — Edited reporting with concise sentences, factual language, and established editorial conventions.
Fiction — Creative writing with dialogue, unusual vocabulary, varied sentence lengths, and more stylistic freedom.
Technical writing — Documentation or explanations that often rely on precise terminology and repetitive structures.
The samples should also be reasonably similar in length where possible. Otherwise, we could accidentally be testing document length rather than genre.
Run Every Sample Through the Same AI Detector
Next, I’d run every document through the same AI detection setup.
Winston AI would be my choice for the main test because it’s specifically designed as an AI detector and provides a straightforward way to check whether content shows characteristics associated with AI-generated writing.
The important part is consistency. Every genre should be tested under the same conditions and the results recorded rather than selecting only the most interesting screenshots.
For each sample, I’d save the genre, word count, source, test date, and Winston AI result.
That gives us something we can actually compare.
What Are We Looking For?
Since every document in this experiment is known to be human-written, the main thing I’d watch for is false positives.
Does one genre get flagged more frequently than another?
For example, technical documentation can contain repetitive terminology because consistency is often desirable. Academic essays can follow predictable structures. Professional emails frequently use standard phrases such as “Thanks for your time” or “Please let me know.”
Fiction, meanwhile, might contain much greater variation in vocabulary and sentence rhythm.
That doesn’t automatically mean an AI detector will treat these genres differently. The point of the experiment is to measure whether it actually happens.
Why Genre Could Matter
AI detectors analyze characteristics within text rather than knowing the actual history of how a document was created.
Different genres naturally produce different linguistic patterns.
Technical writers intentionally repeat terminology for clarity. Journalists may follow established style conventions. Academic writers often use formal transitions. Email writers rely on familiar phrases. Fiction writers may deliberately break grammatical or stylistic conventions.
If a detector reacts differently to those patterns, genre becomes an important variable when interpreting its results.
This is also why I wouldn’t assume that a detector performing well on essays will produce identical performance on technical documentation or fiction.
Don't Test Just One Example
One of the easiest mistakes with AI detector experiments is testing five documents—one from each genre—and treating the results as definitive.
That sample is far too small.
A better experiment would include multiple documents and multiple authors within every category. Ideally, the samples would also vary in writing style while maintaining verified human provenance.
If Winston AI identifies nearly all of them as human, that’s useful information. If certain genres produce more unexpected classifications, that’s useful too.
Either result should be published.
What Would Make the Test More Transparent?
If I were publishing the full experiment on DEV.TO, I’d include the methodology alongside the results.
Readers should know where the human samples came from, how their provenance was verified, their approximate lengths, when the detector was tested, and whether any editing was performed before scanning.
Most importantly, I’d publish the complete results rather than only examples that support a predetermined conclusion.
That makes it possible for someone else to repeat the experiment with Winston AI or compare the same dataset against another AI detector.
What Could This Tell Us About AI Detection?
The goal isn’t to prove that AI detectors work or don’t work based on a handful of documents.
It’s to understand where detection results may change.
If human-written essays, emails, journalism, fiction, and technical documents consistently receive similar classifications, that strengthens confidence across those genres. If results vary significantly, that tells us genre should be considered when interpreting an AI detection score.
For me, Winston AI would be a strong starting point for this kind of controlled AI detector human-writing test. But regardless of the detector used, the methodology matters just as much as the final percentages.
The interesting question isn’t simply:
“Can an AI detector recognize human writing?”
It’s:
“Can it recognize human writing consistently when humans write in completely different ways?”
Top comments (0)