DEV Community

Cover image for Mind the layers: a three-layer model for document AI
Janos Tolgyesi
Janos Tolgyesi

Posted on Originally published at mrtj.pro

Mind the layers: a three-layer model for document AI

In case you missed the previous episode: I argued that there is no single universal schema for what a document means, because what matters about a document is relative to the question you are asking. Knowledge is structured by use, the categories we extract are tools rather than universal truths, and the design move that follows is to equip the inquiry rather than model the document. That was the stance. This post is the architecture.

This is the second of a three-part series. Part one argued the case; part three will get into implementation. Here, in the middle, I want to make it concrete with a model I have come to rely on: the knowledge you pull out of a document lives on three layers, and the layers differ in how much of the work you can reuse from one document or task to the next. Then I will argue the single rule that keeps such a system debuggable when it breaks in production, which it will: never skip a layer. The model is the tidy part; the rule is the part that earns its keep at three in the morning. Let's dive deeper.

The three layers

Here is the model in one sentence: the knowledge you extract from a document sits on three layers, and each layer has both a different reuse profile (how much of it carries over to the next document or task) and a different epistemic status (the kind of question it answers). I will take each layer in three beats: what it is, how reusable it is, and what kind of question it settles.

Layer 1 is the document's intrinsic structure: pages, blocks, tables, reading order, sections, signatures, the geometry of the thing on the page. This layer is fully reusable, because every document is, first and before anything else, a document: a contract, an invoice, a vaccination certificate, and a novel are all pages with blocks and a reading order. Its epistemic status is perception: it answers "what is physically on the page?" It is domain-free and, increasingly, engine-detectable by general-purpose document layout analysis. Build it once, reuse it everywhere.

Layer 2 is the domain's entities and relations: the parties, dates, amounts, issuing authorities, and cross-references that a domain keeps caring about across documents. This layer is partially reusable. A sparse upper ontology generalizes (most legal documents have parties and effective dates), and each domain extends it with its own. Its epistemic status is grounding: it answers "what is this document about, in domain terms?" It is the layer that turns a block of text at coordinates on page four into "the governing-law clause naming the State of New York."

Layer 3 is workflow-specific knowledge: "is this a duplicate payment?", "is this clause enforceable?", "summarize this filing for the board." This layer is not reusable, and it should not be. Its epistemic status is inference: it answers "what does this particular workflow need to conclude?" This is part one's "equip the inquiry" layer, wearing its architectural hat. Each workflow builds the slice of knowledge its question needs, and retires it when the question changes.

None of this is specific to legal or financial documents. Run the same model over a novel and the layers still fall out cleanly: chapters, paragraphs, and dialogue at Layer 1; characters, places, and events at Layer 2; a question like "how does John change across the book?" at Layer 3. I will come back to that.

The following diagram shows the model, and the shortcut it warns against.

Flowchart: a raw document flows down through Layer 1 (intrinsic structure, perception), Layer 2 (domain entities and relations, grounding) and Layer 3 (workflow knowledge, inference) to the workflow answer, while a dashed arrow skips straight from the document to the answer, labelled: errors compound, failure undiagnosable.
Diagram by the author

Hold on to the three epistemic statuses, because they turn out to be the point: perception, grounding, and inference are three different kinds of question, and a system that keeps them apart is a system that can tell you which one it got wrong.

The shapes of Layer 2

Layers 1 and 3 are easy to picture. Layer 2 is the interesting one, because it looks quite different from one domain to the next, and it is where most of the design judgment goes. Three contrasting examples make the point.

The first is the humble invoice, the same example that closed part one. An invoice has a rich, stable Layer-2 vocabulary: issuer, recipient, line items, amounts, tax, dates, a reference number. This is the reusable core I promised to give a home to, and here it is: capture that vocabulary once and reuse it across every workflow that touches invoices. For the invoice, Layer 2 is thick, and it is mostly a matter of naming the entities that are reliably there.

The second is a legal contract, and it behaves differently. Its stable vocabulary is comparatively thin: parties, governing law, a few effective dates. The real Layer-2 work is not naming entities; it is two harder jobs. The first is reference resolution. A contract cites statutes, articles of law, other clauses, and external agreements, and resolving each citation to a stable, canonical identifier (the task known in NLP as entity linking) is what lets an agent reason about the reference rather than merely follow the link: is the cited provision apt, is it still in force, has it been abrogated? Resolving a reference to a canonical identity is the precondition for reasoning about it, because the identifier is a key into a knowledge base where facts about the reference live, not just a pointer to its text. The second is defined-term binding. Contracts define their own terms ("Interest Rate" means what section 1.1 says it means, not what a dictionary says), so binding a defined term to the contract's own definitions section, a within-document coreference between the term and the clause that defines it, is a first-class grounding task. Bind "Interest Rate" to its everyday sense and you have made a grounding error, quietly, and everything Layer 3 concludes from it will be wrong in a way that looks perfectly reasonable.

The third is not a legal or financial document at all, and that is the point of including it. As a thought experiment, take a novel. Its Layer 2 is a cast of characters, the places they move through, the events that occur, and the timeline that orders them. The characteristic Layer-2 work here is neither a fixed vocabulary nor citation resolution but entity tracking across a long narrative: knowing that the "he" on page three hundred is the boy introduced on page five (coreference again), and reconstructing a chronology from a story told out of order. Put a Layer-3 question to it, "how does John's character develop across the book?", and the discipline is unchanged: get the structure and the entities in place before you reason about the arc. The model does not seem to mind the domain. It seems to work.

The point is not that one document is harder than another. It is that Layer 2 is real everywhere, but its thickness varies by domain: for the invoice it is mostly vocabulary, for the contract it is mostly resolution and binding, for the novel it is tracking a cast of entities and a timeline across a long text. The machinery that does the contract job well (detecting references, resolving them, binding defined terms) is a field note of its own, which I will write separately.

Keep Layer 2 sparse, Layer 3 disposable

There is a standing temptation to let Layer 2 grow. A workflow needs some derived fact, computing it once in the shared layer feels efficient, and so a workflow convenience quietly migrates downward into Layer 2. Resist it. Every workflow-specific convenience you bake into Layer 2 erodes the one property that made it worth sharing: its reusability. The discipline is a two-sided rule: keep Layer 2 sparse and stable, and keep Layer 3 rich and disposable.

A concrete case. A due-diligence review and a litigation-risk review run over the same contracts, and they overlap: both care about which obligations survive termination. The instinct is to compute "surviving obligations" once, push it down into Layer 2, and let both share it. Resist it: "surviving obligations" is not a stable domain fact but a task-shaped judgment, and the two reviews define it differently. What belongs in Layer 2 is the raw material they agree on (the termination clause, the obligations it names); the interpretation stays in each review's Layer 3, even though both interpret the same clause. That overlap is not duplication to factor away; hoisting the judgment into Layer 2 would bake one review's assumptions into the shared layer and couple the two. The test: a fact that is task-independent and stable belongs in Layer 2; a judgment whose definition depends on the question stays in Layer 3.

Never skip a layer

Now the rule the whole post is named for. Never go straight from the raw document to the language model. The shortcut is seductive: drop the whole PDF, or a text dump of it, into the prompt, ask your Layer-3 question, and let a capable model sort out the rest. In a demo it works beautifully. In production it fails, and worse, it fails in a way you cannot debug.

The reason is that the three layers are three distinct places where a mistake can happen, and the shortcut collapses them into one. Consider the chain, kept deliberately generic: a table cell is misread (a perception error), so an amount is attached to the wrong party (a grounding error), so the workflow reaches a wrong contractual conclusion (an inference error). That errors cascade and compound down a chain of dependent stages is not a new observation: it is a long-studied failure mode of language-processing pipelines (Finkel, Manning, and Ng, 2006), whose own remedy was to reduce the propagation with joint inference. The layered response makes a different and more modest bet. When you build the layers explicitly, each of those is a separate, inspectable step. When you throw everything into one prompt, all three failure modes happen inside the same opaque call, and when the answer comes back wrong you have no way to ask the only question that matters: which layer failed? Was the number misread, or read correctly and misattributed, or read and attributed correctly but reasoned about wrongly? One prompt cannot tell you. Three layers can.

That is the payoff, and it is worth stating plainly, because it is the entire reason to pay the tax of building the layers at all: explicit layers make a failure localizable to one layer. A layered system fails in a place; a skipped-layer system just fails. And what fails in a place can be tested in a place: each layer takes its own golden dataset, which is far less noisy than one end-to-end metric whose definition of success is hard to pin down (a topic I will take up in its own post). Diagnosability is not a nice-to-have you add later; it is a property you either designed in or gave away on day one. If this sounds like a familiar engineering discipline, it is: separation of concerns applied to document knowledge, the same reason strict layering keeps systems like the OSI stack tractable, where each layer builds on the one below without reaching around it.

Going back to the source is not the same as skipping a layer. A Layer-3 task will often need a detail the generalist Layer 2 never captured, and the right move is to reach back and re-read the exact span the lower layers point to: the full text of clause 4.2, or the chapter where the two characters first meet. That is using the layers to navigate back to grounded evidence, not bypassing them. What the rule forbids is the ungrounded shortcut, asking the final question of the raw document with no layers built at all.

There is a precondition hiding in here, which I will only gesture at now. Layering pays off only if the lower layers hand the upper layers something stable to point at. If Layer 1's identifiers churn every time you re-extract the document (re-run the extraction pipeline after, say, an OCR or model upgrade), then Layer 2's groundings and Layer 3's conclusions dangle, and the whole structure collapses back into re-extract-everything. Stable identity across re-extraction is what makes layering hold together, and it is the subject of part three.

A familiar split

If some of this feels familiar, it may be because you have met it under other names. In a talk I gave on document AI at AWS, I split the pipeline into extraction, knowledge, and generation. Those are the same three layers wearing different labels: extraction is Layer 1, knowledge is Layer 2, generation is Layer 3. I find the layer framing more useful now precisely because it makes the never-skip rule obvious, and the rule is what tells you the split is worth keeping even in an era when a single large model could, in principle, do all three jobs in one call. Being able to skip the layers is exactly why you need a rule not to.

What to carry forward

In summary:

  • Three layers, three reuse profiles. Intrinsic structure is fully reusable, domain entities are partially reusable, and workflow knowledge is not reusable at all (and that last one is a feature, not a gap).
  • Each layer answers a different kind of question: perception, grounding, inference. Keeping them separate is what lets a system tell you which one it got wrong.
  • Never skip a layer. The raw-document-to-model shortcut compounds errors into one undiagnosable failure; explicit layers buy you the ability to say where it broke.

The next time you are tempted to pipe a document straight into a prompt and ask the hard question, add the two intermediate stops: extract the structure, ground the entities, and only then infer. It costs more to build and it pays for itself the first time something goes wrong and you can point at the layer that did it.

The rule leans on a precondition I have only hinted at: the lower layers have to give the upper layers something stable to annotate. In part three I will build that out: a document object model that survives re-extraction, so identity does not churn underneath your groundings and conclusions. Stay tuned.


Janos Tolgyesi is an engineer building document AI systems, currently in legal AI, and an AWS Community Builder since 2020. Field notes on knowledge management, agents, and context engineering in document AI: mrtj.pro.

Top comments (0)