DEV Community

Arthur031221
Arthur031221

Posted on

I tested my Docling JSON checker against real documents and found my own bug

docling-guard command line demo

Docling turns a PDF, DOCX, or scanned page into a structured JSON document. Two open issues against the project caught my attention: one where a text item's source span runs past the end of its own text after dehyphenation, another where a table cell's content bleeds into the column next to it. In both cases the conversion still returns a valid JSON file, so nothing downstream raises an error. A search index or a RAG pipeline just gets a quietly wrong chunk.

What it does

docling-guard is a small Python CLI with no dependencies. It reads a Docling JSON export and checks two things: that every provenance span fits inside the text it claims to describe, and that table cells do not collide or overlap on the page grid. A second command, compare, diffs two exports and reports any drop in text or table count after a Docling upgrade. It does not run OCR or load model weights, it only inspects the JSON Docling already produced.

Testing it for real

The first release shipped with one synthetic fixture, a short document built by hand to demonstrate the checks. That proves the code runs, not that the checks are useful on real documents.

This week I ran it against the 15 real Docling exports in Docling's own test suite: arXiv papers, a German newspaper interview, a technical handbook, right-to-left documents, and tables. It reported errors in 9 of 15 documents, which looked alarming until I read the findings closely.

Almost all were the same false positive. Docling strips list markers like "b." from a text item's visible text field but keeps the full string, markers included, in a separate orig field. The provenance span is measured against orig, not text. My checker compared it against text, so nearly every enumerated list item in the corpus tripped the check.

Fixing that dropped the count from 9 documents to 3. Those 3 show a different, real pattern: a multi-part text item whose combined span runs 2 characters past its own text, the exact dehyphenation overshoot described in the issue that got me started on this.

What is in the repo now

The three fixtures that caught the bug are committed under tests/data/real/, with source and license noted in SOURCES.md. A script, measure_real_documents.sh, reproduces the full 15-document count. The README now states the real numbers instead of only the synthetic one.

What is rough

It only checks the shape Docling exports today: a texts array and a tables array with the fields this version of Docling uses. A large export, over 64 MiB, is rejected outright rather than streamed. And the table overlap check only compares horizontal coordinates, so it can warn on a legitimate unusual layout.

docling-guard is MIT licensed and free to use.

https://github.com/Arthur031221/docling-guard

Top comments (0)