Extraction is only half of the problem.
A document system can tell you:
$12.4M
But a production system often needs another answer:
Where did that number come from?
I spent the last few months investigating how much of that grounding problem can be solved directly from the document.
That led to TonerHound.
The approach uses PDF structure, character coordinates, OCR, geometry, matching, table relationships and visual page evidence.
The current benchmark result is 72.6179% Word Grounding F1.
The more interesting result, however, was understanding the failures that remained.
This article walks through what worked, what failed, and where document-native grounding starts to hit harder problems.
Top comments (0)