Originally published at htpbe.tech. The version on htpbe.tech stays in sync with the latest detection algorithm — refer to it for the canonical text.
Every month we look at aggregate, anonymized data from checks processed by HTPBE and write up what the structural signals tell us about the state of PDF tampering. No file contents, no personally identifiable information – only the structural and metadata patterns the algorithm uses to classify documents.
This report is about proportions and movement, not raw counts. What share of documents came back flagged, which signals fired more or less often than the month before, which origins shifted, and what the recurring tampering shapes looked like. Those are the numbers that mean something; an absolute file count for a single month is noise by comparison.
August is a quiet month, and the interesting story is in what did not move.
What the Denominator Is, Before Any Number
Every share below is a share of processed checks – completed analyses, not unique documents and not unique users. That denominator includes our own internal self-check runs, owner testing, fixture regressions, and retries; it is not deduped. The same file submitted twice counts twice. Processed volume rose over July, but read that purely as throughput – more analyses ran through the pipeline – not as demand, adoption, or anything about the market.
So this is not a population fraud rate, and no single figure here should be read as one. It is a description of what reached our pipeline and how our pipeline classified it. Two scope points belong up front, before the headline:
- The mix is shaped by how HTPBE is used. A large part of the traffic is API integration testing and people checking documents they already suspect – files pre-selected toward the interesting cases. That lifts the flagged share on its own and has nothing to do with how often documents in the world are altered.
- HTPBE performs structural tamper detection. It reads how a PDF was built and rewritten; it does not verify identity, confirm an issuer’s records, or judge whether the content is true. Nothing here is a KYC or identity check, and a verdict is an input to review, not an adverse decision on its own.
The Shape of the Verdicts
The flagged share held at a little over half of processed checks – indistinguishable from July within month-to-month noise. Last month’s reading – described here as approaching six in ten – and August’s both sit inside the same noise band; the change in wording is not a real move. There is nothing here to explain, and we are not going to invent a cause for a number that did not really change.
The confidence tiers, as an August snapshot:
| Verdict tier | Share of August checks |
|---|---|
| Not returned as modified | just under half |
| Flagged – high confidence | about a third |
| Flagged – certain | about one in five |
That top row is not a clean-bill count. It combines files returned as genuinely intact with files returned inconclusive – where the evidence is insufficient to rule modification in or out – and this report does not break the two apart. Read it as not returned as modified, never as confirmed fine. Read the table as an August snapshot and nothing more. The split between the certain and high-confidence tiers shifted at the month boundary, but that shift is a calibration change, not a severity trend – a reclassification of how confidently the engine records a call took effect right around the start of the window, so July and August sit on two different rulers. Comparing the certain-versus-high split across the two months would measure the ruler, not the documents. We present the two tiers side by side for August only and draw no month-over-month line through them.
A note on the tiers: certain describes how confidently the engine made the modification call. It is never certainty about intent, and never a fraud finding.
Source & Origin Mix
Two origin classes led the August snapshot. Consumer-software exports were the largest class at roughly four in ten of processed checks, with institutional documents next at about three in ten. Read those as an August snapshot, not as a month-over-month trend: an in-window reclassification shifted files between origin buckets this month, so – exactly as with the confidence split above – July and August sit on different rulers, and an origin-mix comparison would measure the ruler, not the documents. You can see the full aggregate breakdown on the PDF statistics page; the month figures here are directional against it.
Easy to misread a big consumer-software slice as a big suspicion signal, so one plain point: an inconclusive result is an honest limit on the evidence, not a red flag. When a file’s origin carries too little structural history to support an intact-or-modified call, the engine returns inconclusive rather than forcing one. A larger inconclusive-prone slice means more files where the evidence is insufficient to rule modification in or out – it does not mean more tampering.
On channel, the API and web paths ran roughly even in August. A newly-tracked anonymous web segment also appears in the data for the first time this month – but it only began partway through August, so what we have is a partial month, not an August-representative rate. That makes any channel trend unreadable this month, so we make none. The web path also still includes our own internal automated traffic – one more reason not to read the channel split as a market signal.
Signals That Moved
The cleanest reads are the ones that lean least on any single verdict – the structural composition of what showed up.
Missing creation dates edged up. The share of processed files arriving with no readable creation timestamp at all moved up to about one in five, from roughly a sixth in July. Qualify it carefully: this is a share of processed checks – shaped by the traffic mix, internal and retry runs included – and it describes what the submitted files contained, not a claim that more documents in the wild lack creation dates. There was also a minor end-of-month change to how we extract that field, but it landed too late to shape a month-wide read. Worth watching; not a real-world authoring shift.
The suppressed slices stayed suppressed. Signed documents, post-signature edits, signature removals, embedded files, and JavaScript were each too thin a base this month to quote a rate, so we keep them qualitative. The one signed-document point worth repeating is the standing one, and it is about structural signing integrity only, not the legal or issuer validity of a signature: a signature valid in the viewer does not guarantee the bytes were never altered, because incremental updates appended after signing fall outside the signed scope. Integrity checked at the structural layer, not the signature-validation layer, is what surfaces that.
Incremental Updates
Incremental updates – content appended to a PDF after its original write – split into two numbers that should not be blurred.
Prevalence – how many of all files carried incremental updates at all – ran about one in ten and was flat versus July. The average revision chain on those files stayed short – only a small number of post-write layers stacked per file.
The flag rate among those files – of the ones carrying incremental updates, how many came back flagged – was nearly all of them in August. Read that strictly as an August snapshot: the base is small, and it is not a trend claim in either direction. The mechanism is unchanged – incremental updates let content be added after the original write, and while legitimate workflows produce them (signature application, annotation, form-fill), those clean cases are a minority of the population that reaches a tamper-detection tool.
Representative Cases
These are composite, anonymized illustrations of the recurring shapes the engine resolved this month – not specific files. Each maps to the structural markers that actually drove the verdict, and each describes a modification, or deliberate restraint – never proven fraud or intent.
The reconstructed statement (verdict: modified). A file reads as one coherent statement, but structurally its pages were assembled from more than one independently produced source rather than issued as a single original. It is reported as an assembled file – never as an untouched original.
The rebuilt document (verdict: modified). A file presented as an untouched institutional original was, structurally, built as a new file rather than issued as the original – a rebuilt construction, not a single issuance. The claim is about the structural shape: the structure records a later rebuild. It is not a claim that a specific prior original has been demonstrated. August broadened coverage of this shape.
The scan edited after capture (verdict: modified). A page that genuinely originated as a scan, but whose embedded image shows signs of having been edited and re-saved after capture, is reported as modified. Restraint runs alongside it: August reduced false flags on legitimate documents that include scanned material, so a genuine document is less likely to draw a modified result on that basis alone.
The signed delivery envelope (verdict: not flagged – restraint). A document delivered inside a national electronic-signature container is not reported as modified merely because of the delivery format. It is still analyzed in full, and the container is never treated as evidence that the document is unmodified. This point concerns structural signing integrity only – not the legal or issuer validity of the signature.
Algorithm Development
August was a light release month by our standards – eight versions shipped, well below July’s heavy release cadence. The work leaned toward false-positive reduction, with two coverage broadenings alongside. Described in outcome terms only:
- Fewer false flags. Several releases narrowed misfires on legitimate documents that were drawing a modified result they should not have. Separately, some results the engine cannot stand behind are now surfaced for human review.
- Two coverage broadenings. Detection of documents reworked after their first production was extended, as was detection of images edited and re-saved after capture.
- Display and description fixes. List and single-document views were aligned so a file shows the same verdict everywhere, and a few public marker descriptions were made clearer – with the underlying detection unchanged.
We are deliberately not drawing a line from this release mix to the flagged share. The in-window releases pushed in both directions, the net is not separable from the shifting intake, and the flagged share did not meaningfully move anyway. Comparability with July is partial, and that is the honest description of it.
PDF Version Landscape
Concentration re-tightened. PDF 1.7 rose to about 38% of the sample, reversing the loosening we noted in July, with PDF 1.4 easing to roughly a quarter. The rest trailed, and PDF 2.0 stayed a rounding-error share despite years of availability. This is directional only – no significance test was run – and, like the missing-creation-date read, it is a detector-independent carrier: it reflects the submission mix this month, not a shift in real-world version adoption.
Summary
August 2026, in relative terms:
- The flagged share held at just over half of processed checks – essentially flat versus July, within noise, and we attribute the non-move to nothing.
- The certain-versus-high split shifted at the month boundary as a calibration change, presented as an August snapshot only; it is not a severity trend and not comparable across the two months.
- Consumer-software origin stayed the largest class at roughly four in ten, institutional next at about three in ten – an August snapshot; an in-window reclassification also shifted the composition.
- Missing creation dates edged up to about one in five as a share of processed files, and incremental-update files held at about one in ten with short revision chains – both carrier reads shaped by the traffic mix, not real-world change.
- PDF 1.7 re-concentrated to about 38% – directional, a submission-mix carrier, not adoption.
Every pattern here comes from the same forensic engine that teams run on their own intake stream through the PDF tamper detection API. If you want to run a single document through the same analysis by hand, the free checker does it in the browser.
Previous report: PDF Integrity Report: July 2026.
This report covers checks processed by HTPBE in August 2026. We analyze only file structure, never document content; web uploads may be retained in anonymized form to improve detection. All figures are aggregate and anonymized.
Top comments (0)