VERIDOC: AI-Powered Document Forensics for Detecting Suspicious Documents
What if a document looks completely legitimate—but someone secretly changed a few details?
A modified font, inconsistent layout, altered metadata, or an unexpected image can sometimes be the clue that something is wrong.
For our hackathon project, we built VERIDOC, an AI-assisted document forensics platform designed to perform a first-level screening of digital documents for signs of tampering.
It does not claim to definitively prove that a document is forged. Instead, it collects multiple forensic signals and uses them to identify documents that deserve further investigation.
Digital documents are everywhere—certificates, reports, applications, invoices, records, and other important documents.
The problem is that editing tools make it increasingly easy to modify these documents while keeping them visually convincing.
A simple visual inspection may not reveal:
- A different font used in one section
- Unexpected formatting changes
- Metadata inconsistencies
- Embedded images or unusual visual elements
- Differences between document versions
- Other subtle structural anomalies
We wanted to build a system that could automatically collect these signals and present them in a way that is easier to understand.
VERIDOC
VERIDOC analyzes an uploaded document through several independent forensic layers.
The overall pipeline looks like this:
Document Upload
↓
Document Ingestion
↓
Text / OCR Analysis Metadata Analysis Visual & Formatting
↓
Evidence Aggregation
↓
Gemma
↓
Risk Assessment
↓
Low / Suspicious / High
↓
Results Deshboard
We used a combination of Python-based document processing, computer vision/forensics techniques, a database layer, and Gemma for AI-assisted reasoning.
- Python
- Streamlit
- SQLite
- Gemma
- PyMuPDF
- OpenCV
- pdfplumber
- Tesseract OCR
- Pillow
We divided the project into four major areas:
- Document Processing & OCR
- Document Forensics & Visual Analysis
- Gemma Integration & Risk Scoring
- Frontend, Dashboard & Integration This allowed us to develop different components independently and bring them together into one system. What We Learned Building VERIDOC showed us that document verification is more than simply reading text. Content + Metadata + Formatting + Visual Evidence + AI Reasoning can provide a much richer picture of a document's authenticity. We also learned the importance of keeping AI grounded in observable evidence instead of asking it to make unsupported decisions.
Links
Final Thoughts
VERIDOC asks three simple questions:
What looks unusual?
What evidence supports it?
How suspicious is the document overall?
That's our approach to making document forensics more accessible and explainable with AI.
Top comments (1)
The "first-level screening, not proof of forgery" boundary is worth carrying into the Low / Suspicious / High result too. A useful paired fixture would be one untouched PDF, the same content legitimately re-exported through another editor (font/metadata changes), and a copy with just the invoice amount changed while preserving the layout.
That tests both false alarms from a benign export and a small edit with few structural clues. I'd want the dashboard to separate "signals found" from "this layer could not be analyzed", especially after OCR failure, so missing evidence doesn't read as a clean bill of health. I haven't run VERIDOC; these are test cases suggested by the pipeline in the article, not a finding about its accuracy.