Originally published at ictlms.net
Universities are quietly removing AI detection percentages from integrity decisions. Not the tools themselves in every case, but their standing as evidence: score it, sure, then treat it as an indicator and go find something real. If your assessment platform's answer to AI cheating is a number from a detector, you're holding the one thing panels have started refusing to act on.
This shift happened faster than most vendors admit. It came from false positives, particularly against students who write in English as a second language, and from faculty who got tired of defending a percentage they couldn't explain. The question now isn't whether you can detect AI text. It's what you can put in front of a panel when a student says they didn't do it.
What actually changed
Two things moved at once, and they reinforce each other.
The first is the detection retreat. A growing set of institutions limited or removed AI scores from misconduct determinations after system-wide reviews. Plenty still run detection, but the instruction to instructors changed: a flag opens a conversation, it doesn't close one.
The second is a shift in what gets assessed. From the 2026 intake, a lot of departments are marking process rather than output. Engineering students get asked to use AI to draft a design and then verify every figure by hand. Literature students submit the essay plus the drafts plus a revision memo explaining what changed and why. Oral defences, in-class writing and staged submissions are all coming back, and even a modest 10 to 20 percent shift of grading weight toward supervised work cuts your exposure to unsupervised submissions considerably.
Both changes point the same way. The artefact alone no longer proves anything. The record around the artefact does.
What a detector score cannot survive
Put yourself on an appeals panel. A student sits in front of you and says they wrote it. What have you got?
A percentage has no working shown. The student can't contest it because there's nothing to contest, the vendor won't come and explain the threshold, and the false-positive literature is now easy enough for any student to find and cite. I've seen a case turn entirely on a student producing a Google Docs version history that the institution's own system hadn't captured. The detector said one thing, the timeline said another, and the timeline won, as it should have.
Everything in the right-hand column is different in one specific way: each line is a fact with a timestamp attached. Panels are perfectly comfortable weighing facts. They are not comfortable weighing a number nobody in the room can account for.
Tiers are replacing bans, and tiers need enforcement
Blanket AI bans have mostly collapsed under their own unenforceability. What's replacing them is a tiered declaration model: no AI for this assessment, limited AI for that one, AI expected and marked on this project. Students declare what they used, often with a short reflection attached.
It's a sensible policy. It's also completely hollow unless the platform underneath can enforce the tier and prove which tier applied.
Look at what each column asks for. Tier 0 needs a locked environment and an identity check that happened before the paper opened. Tier 1 needs draft history kept next to the submission, so the process is visible. Tier 2 needs the prompts and the student's own verification of what came back, because the skill being marked is the checking, not the generating.
Now notice what's missing from all three. There's no detector percentage anywhere in that table, and the tiers work fine without one.
The record your platform should be producing
If I were auditing an assessment platform this term, I'd ask for a single export for one student on one paper and check whether it contains seven things.
Who sat down, verified before the paper opened rather than after. The timing of each answer, started and finished, because unusual timing is the signal that holds up best under scrutiny. Paste events and tab switches with timestamps, not a summary count. Draft or revision history where the assessment allows drafting. The declared AI tier for that paper, stored on the attempt rather than in a course handbook. Every flag the system raised and the name of whoever reviewed it. And the human decision with the reason written down.
That last one gets skipped constantly and it's the one that matters most on appeal. A flag with no recorded reviewer is worse than no flag, because it shows the system noticed something and nobody looked.
The export test is the honest test, by the way. Vendors will tell you all of this is captured. Ask to see it as one file for one attempt. We've written before about why a flag should never be a verdict, and this is the same principle with the paperwork attached.
The identity gap nobody closes
Here's a piece that gets left out of the AI conversation entirely. Every evidence trail above assumes the right person sat the paper. If your identity check is a webcam photo compared by eye at the start, your beautiful timestamped record documents the behaviour of someone you can't name with confidence.
That weakness predates AI and got worse with it, since the same tooling that writes essays now also produces convincing face and voice material. We went through the specifics of that in our piece on identity checks and injection attacks. For the purposes of this post, just treat identity as the foundation the rest of the evidence sits on. A misconduct case built on a shaky identity check falls over at the first question.
Where I'd actually spend the effort
If your budget this year is a single change, make it draft and timing capture rather than better detection. Detection accuracy is a race you can't win and can't defend. Timing and process data is cheap to collect, hard to fake convincingly, and reads clearly to a non-technical panel.
Second priority is the declaration. Getting students to state their tier and, on tiered work, write two sentences about what they used, does more for integrity than any scanner. It moves the burden from catching people to making them account for their own process, and most students will tell you the truth if the form is right in front of them.
Detection tools aren't useless. Keep them if you have them. Just demote them to what they are, which is a prompt to look closer, and never let a percentage be the reason anyone gets a finding against their name.
FAQ
Are universities banning AI detectors outright?
Some have removed the scores from integrity decisions, others still run detection but instruct staff to treat results as indicators only. Complete bans are rarer than the headlines suggest. The consistent change is the score losing its status as sufficient evidence on its own.
If detection is unreliable, how do you catch anyone?
The same way misconduct was caught before detectors existed, with better data. Timing anomalies, work that doesn't match a student's demonstrated ability, drafts that appear fully formed in one paste, and an oral check when something doesn't add up. It's slower and it holds up.
Does a smart online exam need to record keystrokes?
Full keystroke logging is heavier than most institutions want, both technically and for privacy. Paste events, tab switches and per-answer timing give you most of the signal at a fraction of the intrusion. Start there and only go further if your threat model genuinely requires it.
How long should this evidence be kept?
Long enough to cover your appeals window plus any regulatory retention that applies to assessment records, and no longer. Behavioural data about students is exactly the sort of thing you don't want sitting around for five years because nobody set a policy.
What about students who genuinely write like an AI?
They exist, and they're disproportionately people writing in a second language or students who've been taught a formulaic structure. That's precisely why the score can't stand alone. The evidence record protects those students as much as it catches anyone.
Does ICTExam capture this?
Per-answer timing, identity verification before the paper opens, paste and tab events, flags with a named reviewer and a recorded human decision, yes. Draft history depends on how the assessment is configured. The features page lists what's in the current release, and we'd rather you test the export on a real attempt than take our word for it.
Try the export test on whatever platform you run this term. One student, one paper, one file. What comes back will tell you whether your integrity process has evidence behind it or just a number.
Top comments (0)