GPTHuman's Own Scores Cannot Verify Facts or Citations. Here's a Deterministic Test
This week I opened GPTHuman's output-quality page to find out what its three scores were supposed to settle. I copied Human Score, Similarity Score, and Readability Score into my notes, then tried to add a fourth column: did the evidence survive?
That column stayed blank.
GPTHuman's own output-quality guide says the three numbers describe different features of the result and that none is a complete quality measure. It also says they do not independently verify factual accuracy, quotations, citations, plagiarism, policy compliance, authorship, or the verdict of an outside detector.
That is a useful limitation, not a footnote to skip. It tells you what kind of test to build.
The Human Score concerns patterns that GPTHuman's system associates with human-like writing. Similarity concerns how closely the result resembles the source. Readability uses the Flesch-Kincaid method to estimate how easy the text is to read. A result can look strong on all three and still change not to nothing, move a citation to the wrong claim, or reverse a threshold.
So I would not use any of those scores as the acceptance criterion. I would give the tool a source with facts that have only one correct answer, then mark each answer pass or fail.
Here is the source. It is 83 whitespace-delimited tokens. The name, event, and figures are invented, so no real paper or private data needs to enter the test.
Rivera (2025, p. 12) wrote, "Batch A stayed below 4.75 ms." The audit covered 128 records on 12 March 2025. Batch A contained 64 records and 9 exceptions; Batch B contained 64 records and 8 exceptions. The combined exception rate was 13.28% (17/128). The rate applies only when room temperature is at or below 21.5 C. Because the sensor failed after record 96, records 97-128 were excluded from the latency average but remained in the exception count. The conclusion is preliminary, not final.
Put this table directly under it:
| Cohort | Records | Exceptions | Rate |
|---|---|---|---|
| Batch A | 64 | 9 | 14.06% |
| Batch B | 64 | 8 | 12.50% |
| All | 128 | 17 | 13.28% |
The paragraph is short on purpose. GPTHuman's meaning-preservation guide warns that exact preservation cannot be assumed in every passage. It specifically tells users to check names, dates, numbers, units, quotations, citations, conditions, exceptions, and qualifications. This fixture puts most of those failure points into a space you can inspect in a few minutes.
Here is the ledger I would use:
| Check | Fixed baseline | Pass condition |
|---|---|---|
| Citation |
Rivera (2025, p. 12) belongs to the sentence that follows |
Author, year, locator, and attachment stay unchanged |
| Quotation | "Batch A stayed below 4.75 ms." |
Quotation marks and every word inside them stay unchanged |
| Dates |
2025 and 12 March 2025
|
Both dates keep their value and role |
| Units |
4.75 ms and 21.5 C
|
Values remain attached to the same units |
| Arithmetic | 9/64 = 14.06%, 8/64 = 12.50%, 17/128 = 13.28% | All three rates still recalculate after rounding to two decimals |
| Condition | The rate applies only at or below 21.5 C |
only, the downward direction, and the inclusive 21.5 C boundary all remain |
| Exclusion | Records 97-128 leave the latency average but stay in the exception count | The two populations are not merged, swapped, or shortened |
| Qualification | The conclusion is preliminary, not final | Uncertainty is not promoted into certainty |
| Table | Four headers, three rows, twelve body cells, fixed row order | Every label, value, and position survives |
This is intentionally unforgiving. If the output changes 13.28% to 13.3%, that may be acceptable for some jobs, but it fails this fixture because the baseline asked for two decimal places. If at or below becomes below, the new sentence excludes exactly 21.5 C. That fails too. Fluent prose does not get to overrule the source.
Do not merge the nine rows into a single percentage. A pass rate of eight out of nine hides whether the failed row was a table label or the word not. Keep the failure attached to the thing that changed.
The setup around the run matters as much as the text. Save an untouched copy first. Record the language, mode, plan, date, and whether the source was pasted or imported. Change one setting at a time. Save each result under a new filename.
Then compare exact tokens before judging style. Recalculate the three rates. Read the condition, exclusion, and qualification as logic, not as approximate meaning. Count the table cells after copying the result out of the browser.
GPTHuman's testing guidance makes the same separation at a broader level. It says meaning preservation, factual consistency, readability, language quality, originality, and named-detector results answer different questions. The page does not publish a universal independently verified bypass rate, and it says a defensible detector report needs a defined sample set, settings, detector versions, date, threshold, and complete outcomes.
I did not run GPTHuman or a detector for this article. There is no pass rate to report. The point is to define the test before seeing the output, so a pleasant result cannot quietly move the goalposts.
One successful run would still support only a narrow claim: this fixture passed under this language, mode, input route, plan, and date. It would not establish that every citation survives, every subject behaves the same way, or a future detector will agree.
GPTHuman's responsible-use page also tells users to verify facts, calculations, quotations, and citations and to follow the rules that apply to the work. That last part stays outside the ledger. A technically preserved result can still be disallowed by a course, employer, publisher, or client.
The three scores can help you inspect an output. They cannot approve it. For evidence-heavy writing, the boring little ledger is the more useful result.
Originally published at HumanPen: https://humanpen.net/blog/gpthuman-ai-humanizer-review?utm_source=devto&utm_medium=article&utm_campaign=d163
Top comments (0)