<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Felipe Cardona</title>
    <description>The latest articles on DEV Community by Felipe Cardona (@felipe_anyformat).</description>
    <link>https://dev.to/felipe_anyformat</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4114113%2Fc529dc1a-df38-484a-aef0-0d954581ce7e.png</url>
      <title>DEV Community: Felipe Cardona</title>
      <link>https://dev.to/felipe_anyformat</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/felipe_anyformat"/>
    <language>en</language>
    <item>
      <title>LlamaParse vs Unstructured vs Reducto: Which One Fits Your RAG Pipeline?</title>
      <dc:creator>Felipe Cardona</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:34:34 +0000</pubDate>
      <link>https://dev.to/felipe_anyformat/llamaparse-vs-unstructured-vs-reducto-which-one-fits-your-rag-pipeline-55io</link>
      <guid>https://dev.to/felipe_anyformat/llamaparse-vs-unstructured-vs-reducto-which-one-fits-your-rag-pipeline-55io</guid>
      <description>&lt;p&gt;LlamaParse, Unstructured and Reducto are the three tools that come up first when a team is choosing how to get documents into a RAG pipeline or an LLM-ready format. All three are AI-native, all three have real engineering behind them, and all three keep shipping fast enough that a comparison from even three months ago is out of date. This one is current as of September 2026.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Method.&lt;/strong&gt; Facts below come from each vendor's own documentation and pricing pages, checked in September 2026, with every accuracy number attributed to whoever published it. None of these three vendors publish numbers on a shared, independent benchmark against each other today, so this piece does not force one. Where a number exists, it says whose it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;LlamaParse&lt;/th&gt;
&lt;th&gt;Unstructured&lt;/th&gt;
&lt;th&gt;Reducto&lt;/th&gt;
&lt;th&gt;anyformat&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core job&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Parse documents into LLM-ready Markdown/JSON for RAG&lt;/td&gt;
&lt;td&gt;Chunk documents into element arrays for vector stores; now also schema extraction&lt;/td&gt;
&lt;td&gt;Parse and extract via API, code-first&lt;/td&gt;
&lt;td&gt;Schema-defined extraction into business systems, with review and evaluation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Structured field extraction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Extract, Classify, Split added 2026)&lt;/td&gt;
&lt;td&gt;Yes (Extract node, JSON schema)&lt;/td&gt;
&lt;td&gt;Yes (code-defined schemas)&lt;/td&gt;
&lt;td&gt;Yes (zero-shot, no-code schema)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Per-field confidence&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claimed on 3 of 5 tiers, self-reported, no published eval&lt;/td&gt;
&lt;td&gt;Not offered&lt;/td&gt;
&lt;td&gt;Per-field confidence, no calibration claim&lt;/td&gt;
&lt;td&gt;Calibrated, with a visual citation on every field&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Human review queue&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No (confidence/citations sit behind an &lt;code&gt;expand&lt;/code&gt; call)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes, corrections flow back into the run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Evaluation against your own ground truth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not offered&lt;/td&gt;
&lt;td&gt;Not offered&lt;/td&gt;
&lt;td&gt;Not offered&lt;/td&gt;
&lt;td&gt;Yes (Health tab, numbered/immutable runs)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;On-premise / air-gapped&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Enterprise self-hosted / VPC&lt;/td&gt;
&lt;td&gt;Self-hosted (dedicated instance, VPC, bare metal)&lt;/td&gt;
&lt;td&gt;Yes, incl. air-gapped&lt;/td&gt;
&lt;td&gt;Yes, incl. air-gapped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;EU jurisdiction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;US company; EU region live since July 2026&lt;/td&gt;
&lt;td&gt;US company; FedRAMP High, ISO 27001, SOC 2, GDPR-compliant&lt;/td&gt;
&lt;td&gt;US company; EU regional endpoints&lt;/td&gt;
&lt;td&gt;EU-native&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pricing model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Credits; parse $1.25–$56.25/1,000 pages by tier, plus a per-form surcharge&lt;/td&gt;
&lt;td&gt;Free 10k pages, then $0.015/page ($15/1,000); custom Business tier&lt;/td&gt;
&lt;td&gt;Flat $10/1,000 pages on &lt;code&gt;r-1&lt;/code&gt; (preview); legacy $15–30/1,000&lt;/td&gt;
&lt;td&gt;Credits per operator: Parse ~€25/1,000, Extract ~€35/1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  LlamaParse
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.llamaindex.ai/llamaparse" rel="noopener noreferrer"&gt;LlamaParse&lt;/a&gt; is LlamaIndex's document platform, built to turn documents into LLM-ready Markdown and, since 2026, extended with Extract, Classify and Split on the same credit wallet. Its biggest practical draw for a RAG team is distribution: it's a listed connector in Claude's directory and publishes a European MCP endpoint, so an agent can start using it with very little setup. Pricing spans a wide range by tier, from $1.25 to $56.25 per 1,000 pages, plus a 10-credit surcharge added in September 2026 for any page containing a form, so the real invoice depends heavily on document mix. Confidence scores left beta with a stated calibration on three of its five tiers (LlamaParse's own number: at a 0.8 threshold, roughly 75% of extraction errors fall below the line), but nothing routes a low-confidence field to a person; confidence and citations sit behind an &lt;code&gt;expand&lt;/code&gt; call on the API response, not a review interface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; teams already in the LlamaIndex ecosystem who want RAG-ready output fast and are comfortable resolving low-confidence fields themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unstructured
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://unstructured.io/" rel="noopener noreferrer"&gt;Unstructured&lt;/a&gt; is a partially open-source platform built for RAG ingestion, with the widest connector ecosystem of the three (71+, including Databricks, Elasticsearch, S3 and Google Drive). It has grown past pure chunking: its Extract node now takes a JSON schema and returns extracted values inside the same ingestion pipeline, so "does it do structured extraction" is no longer the dividing question it used to be. What it still doesn't return is a per-field confidence score, so an uncertain value has nowhere to route. Its commercial platform holds SOC 2, HIPAA, ISO 27001 and, more recently, FedRAMP High authorization (&lt;a href="https://unstructured.io/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;); the open-source library isn't in that scope. Unstructured publishes its own SCORE benchmark (0.917 Adjusted CCT, 0.027 hallucination rate), and an independent Procycons benchmark from March 2025 measured it at 51 seconds per page, well behind Docling and LlamaParse on speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; teams that need the broadest connector coverage for vector-store ingestion, and can now also lean on it for schema extraction if per-field confidence isn't a requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reducto
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://reducto.ai/" rel="noopener noreferrer"&gt;Reducto&lt;/a&gt; replaced its multi-pass OCR-and-VLM pipeline in September 2026 with &lt;code&gt;r-1&lt;/code&gt;, a single end-to-end model priced flat at $10 per 1,000 pages (&lt;a href="https://docs.reducto.ai/reference/credit-usage" rel="noopener noreferrer"&gt;pricing reference&lt;/a&gt;), currently in preview and API-only. It publishes the open RD-TableBench table benchmark and reports strong numbers there, and its product suite also includes an Edit endpoint for filling and writing back to documents (anyformat has an equivalent Edit node too, currently in Beta). What Reducto doesn't ship is a human review loop: extracted values, including &lt;code&gt;r-1&lt;/code&gt;'s new bounding-box citations, come back with no reviewer surface to act on them, and citations disable chunking on their side. Version pinning exists at the model level, but on a 4-week deprecation clock rather than a permanent guarantee. It holds SOC 2 and HIPAA, offers EU regional endpoints, but doesn't list ISO 27001.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; engineering teams that want a fast, cheap parsing primitive and are building the review and orchestration layer themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where anyformat fits
&lt;/h2&gt;

&lt;p&gt;anyformat isn't trying to win the RAG-ingestion conversation these three are having; the pipelines it's built for end in a business system, not a vector store. The overlap is real, though: all four now do schema-defined structured extraction. The difference is what comes back with the value and what happens after. Every field carries a calibrated confidence score and a visual citation to where it was read; low-confidence fields route to a named reviewer, and the correction is recorded against the run. Every Extract and Classify workflow has a Health tab: build a dataset from verified ground truth, run numbered, immutable evaluations against any workflow version, and compare the result to live production accuracy before you ship a change. None of the other three in this comparison have that measurement loop today.&lt;/p&gt;

&lt;p&gt;anyformat's own published accuracy numbers come from two different sources, and it's worth keeping them separate: the landing-page Parse Score benchmark (78.1% at $25 per 1,000 pages) measures cost-adjusted quality against frontier LLMs, not against these three vendors. On an anyformat-run internal harness (South Summit), anyformat scored 90.7 against LlamaParse at 85.3 and Reducto at 77.0; that number is anyformat-published, on anyformat's own harness, not an independent result, and Unstructured wasn't part of that run. Treat it as a data point to verify on your own documents, the same way this piece treats every other vendor's self-published number.&lt;/p&gt;

&lt;p&gt;Pricing is credit-based per operator: Parse runs 25 credits per page (roughly €25 per 1,000 pages), and schema extraction adds Extract at 35 credits per page (roughly €35 per 1,000 pages) on top, so the full extraction pipeline runs about €60 per 1,000 pages, with calibrated confidence, review and evaluation included rather than billed as separate add-ons. Deployment covers cloud, private cloud and on-premise including air-gapped environments, with ISO 27001 certification and zero-retention processing by default, and anyformat is EU-native rather than a US company with EU-region options.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; teams whose pipeline output has to be trusted enough that a named person signs off on it, not just parsed well enough to embed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to choose
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You're building a RAG index and want the fastest path to Markdown.&lt;/strong&gt; LlamaParse's tiering and Claude-directory listing make it the easiest to bolt onto an agent quickly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need the widest connector coverage for ingestion.&lt;/strong&gt; Unstructured's 71+ connectors are unmatched here, and its Extract node now covers basic schema extraction too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You want a cheap, fast parsing primitive and will build the rest yourself.&lt;/strong&gt; Reducto's &lt;code&gt;r-1&lt;/code&gt; is hard to beat on raw parse price; budget separately for review, orchestration and evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The extracted values go into a system someone is accountable for.&lt;/strong&gt; That's the anyformat case: confidence you can set a threshold on, a review queue that records what changed, and evaluations against your own ground truth before a workflow update ships.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;EU jurisdiction is a hard requirement.&lt;/strong&gt; anyformat is the only EU-native platform of the four; the other three offer EU regions or certifications, not EU governance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which is more accurate: LlamaParse, Unstructured or Reducto?&lt;/strong&gt;&lt;br&gt;
There's no independent benchmark that scores all three the same way today, so any single answer is misleading. Each publishes its own numbers on its own benchmark (LlamaParse's ParseBench, Reducto's RD-TableBench, Unstructured's SCORE), and where they do compare each other directly, the numbers disagree sharply: Reducto's own &lt;a href="https://reducto.ai/compare/reducto-vs-llamaparse" rel="noopener noreferrer"&gt;"LongExtractionBench"&lt;/a&gt; claims 99.6% precision and recall against LlamaParse's 80.0% and 77.5%, a benchmark Reducto commissioned and LlamaParse hasn't validated. Test on your own documents before treating any vendor's number, including a rival's number about another rival, as decisive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can Unstructured and Reducto do structured data extraction now?&lt;/strong&gt;&lt;br&gt;
Yes, both added it. Unstructured's Extract node takes a JSON schema; Reducto's Extract endpoint is code-defined. Unstructured doesn't return a per-field confidence score with the extracted value at all. Reducto does return per-field confidence, but without a published calibration claim, so a 0.9 isn't guaranteed to mean the same thing from one field to the next. Neither ships a reviewer surface to route an uncertain field to, calibrated or not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is anyformat a replacement for LlamaParse or Unstructured?&lt;/strong&gt;&lt;br&gt;
Not for RAG ingestion specifically; that's a different job. Some teams run both: one of these three for vector-store ingestion, and anyformat for the documents whose extracted fields have to be reviewed, corrected and evaluated before they reach a business system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which of the three is cheapest?&lt;/strong&gt;&lt;br&gt;
On list price for parsing alone, Reducto's &lt;code&gt;r-1&lt;/code&gt; at a flat $10 per 1,000 pages is the cheapest today, followed by Unstructured's pay-as-you-go tier at $15 per 1,000. LlamaParse's range depends heavily on tier and document mix. None of these figures include a review or evaluation layer, since none of the three ship one.&lt;/p&gt;




&lt;p&gt;If your pipeline needs to end in a decision someone signs off on rather than a vector store, see how anyformat compares directly: &lt;a href="https://anyformat.ai/vs/anyformat-vs-reducto" rel="noopener noreferrer"&gt;anyformat vs Reducto&lt;/a&gt;, &lt;a href="https://anyformat.ai/vs/anyformat-vs-unstructured" rel="noopener noreferrer"&gt;anyformat vs Unstructured&lt;/a&gt; and &lt;a href="https://anyformat.ai/vs/anyformat-vs-llamaparse" rel="noopener noreferrer"&gt;anyformat vs LlamaParse&lt;/a&gt; each go field by field.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://anyformat.ai/blog/llamaparse-vs-unstructured-vs-reducto" rel="noopener noreferrer"&gt;anyformat.ai&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>rag</category>
      <category>llm</category>
      <category>ai</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>How to reduce LLM hallucinations in document extraction</title>
      <dc:creator>Felipe Cardona</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:27:17 +0000</pubDate>
      <link>https://dev.to/felipe_anyformat/how-to-reduce-llm-hallucinations-in-document-extraction-ppp</link>
      <guid>https://dev.to/felipe_anyformat/how-to-reduce-llm-hallucinations-in-document-extraction-ppp</guid>
      <description>&lt;p&gt;A hallucination in LLM document extraction is a value the model returns that does not appear in the document: an invented invoice total, a date lifted from the wrong field, a supplier name completed from the model's memory instead of the page. In an extraction pipeline this is worse than a missing value. A missing value gets flagged. A hallucinated one gets paid.&lt;/p&gt;

&lt;p&gt;Hallucinations in this context tend to fall into a few recognizable patterns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fabrication&lt;/strong&gt; — a value with no basis anywhere in the document (an invented total).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Misattribution&lt;/strong&gt; — a real value from the document, pulled onto the wrong field or row (a date lifted from an unrelated field).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory completion&lt;/strong&gt; — a plausible value the model already "knows" from training, substituted for what the page actually says (a supplier name filled in instead of read off the page).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Format hallucination&lt;/strong&gt; — an invented value shaped closely enough like a real one (a well-formed date, a valid-looking ID) to pass a casual review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The seven techniques below reduce how often that happens. The more important shift is conceptual, so it goes first: you cannot bring hallucinations to zero, but you can make every one of them detectable, and a detectable hallucination is a review task with a cost, no longer a risk you carry blind. That principle separates production extraction systems from demos.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Technique&lt;/th&gt;
&lt;th&gt;What it catches&lt;/th&gt;
&lt;th&gt;What it does not catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Grounding&lt;/td&gt;
&lt;td&gt;Values with no location in the document&lt;/td&gt;
&lt;td&gt;A wrong value that does exist elsewhere on the page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Calibrated confidence&lt;/td&gt;
&lt;td&gt;Uncertain extractions, flagged for review&lt;/td&gt;
&lt;td&gt;Confident errors (these need 6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Schema constraints&lt;/td&gt;
&lt;td&gt;Wrong types and shapes (prose in a date field)&lt;/td&gt;
&lt;td&gt;Plausible values of the right type&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Layout-aware parsing&lt;/td&gt;
&lt;td&gt;Errors caused by scrambled reading order and broken tables&lt;/td&gt;
&lt;td&gt;Errors in clean text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5. Threshold routing&lt;/td&gt;
&lt;td&gt;Turns 2 into an operating decision&lt;/td&gt;
&lt;td&gt;Nothing by itself; it is the control loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6. Deterministic validation&lt;/td&gt;
&lt;td&gt;Confident errors that break arithmetic or checksums&lt;/td&gt;
&lt;td&gt;Errors in free-text fields&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7. Continuous measurement&lt;/td&gt;
&lt;td&gt;Drift after model or prompt changes&lt;/td&gt;
&lt;td&gt;Nothing in real time; it is the feedback loop&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  1. Ground every extracted value in the source document
&lt;/h2&gt;

&lt;p&gt;Grounding means every extracted field carries a pointer to the exact place in the document where the value was read: page, text span and, where the system supports it, the region on the page. If a value cannot be traced to the page, it should not survive the pipeline. Grounding turns "trust the model" into "click and check": a reviewer verifies a flagged invoice total in seconds because the extraction shows where it came from.&lt;/p&gt;

&lt;p&gt;In anyformat, every extracted field is returned with its evidence in the same API response: the source text and the page number it was read from. In the review interface the evidence is highlighted on the page itself, so the person checking a field sees the location, not just the text. It also works in reverse: clicking a value on the page surfaces the field it was extracted into, so a reviewer can start from either the extraction or the document and land on the same place.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Use calibrated confidence scores, not raw model confidence
&lt;/h2&gt;

&lt;p&gt;A confidence score is the number an extraction system attaches to a value to say how sure it is that the value is correct; it is calibrated when that number matches empirical accuracy, so a field scored 95 is correct 95% of the time, measured on real documents. Raw token probabilities from an LLM are not calibrated, and they tend to be highest exactly where the model is filling a gap. Calibration means evaluating the extraction system against ground truth on a large document set and adjusting the scores until they mean what they say.&lt;/p&gt;

&lt;p&gt;anyformat's per-field confidence is calibrated: the score is derived from the extraction model's token probabilities and fitted against ground truth (Platt scaling), so a 90 means the same thing on a supplier invoice as on a bank statement.&lt;/p&gt;

&lt;p&gt;Confidence is measured at each stage separately, not just once at the end: parsing has its own score for how well the document's structure and text were read, extraction has its own per-field score, and a document-level rollup summarizes both. That separation is what lets you tell whether a bad result came from a parsing problem or an extraction problem, instead of one opaque number covering the whole pipeline.&lt;/p&gt;

&lt;p&gt;Two things this page does not claim: that calibration is unique to anyformat, and that any vendor's calibration is proven by the word alone. The evidence for a calibration claim is a published evaluation, which is what the number above is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Constrain extraction with an explicit schema
&lt;/h2&gt;

&lt;p&gt;Schema-constrained extraction forces the model to return values of a declared type and shape: a date field cannot return prose, a currency field cannot return a sentence. A field typed as a date either returns a valid date like &lt;code&gt;2026-03-15&lt;/code&gt;, or nothing: it cannot return &lt;code&gt;"sometime in March"&lt;/code&gt;, or a stray &lt;code&gt;"xxx"&lt;/code&gt; the model produced when the page was unclear. Free-form extraction invites the model to write. A schema forces it to read. Zero-shot schema definition means the same schema works on layouts the system has never seen, without training a template per supplier. Schemas do not stop a plausible wrong value of the right type; that is what techniques 2 and 6 are for.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Parse the layout before extracting: reading order is a hallucination source
&lt;/h2&gt;

&lt;p&gt;Many extraction hallucinations are parsing failures wearing a different name: merged table cells, multi-column reading order, headers attached to the wrong section. If the parser feeds the model scrambled text, the model fills the gaps by inventing. Layout-aware parsing (tables reconstructed as tables, reading order preserved, page structure explicit) removes the ambiguity the model would otherwise hallucinate into. The effect is largest on complex tables and on long documents, where context degradation compounds with every page.&lt;/p&gt;

&lt;p&gt;anyformat publishes its parsing quality: a Parse Score of 78.1% on a set of more than 1,000 real documents, ahead of Gemini 3.5 Flash (77.9%) and GPT-5.6 (74.2%) on the same table. On ParseBench, LlamaIndex's independent benchmark of about 2,000 human-verified pages, anyformat scores 80.83, ranking 4th of the 35 engines evaluated.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Route by confidence threshold: automate the confident, review the uncertain
&lt;/h2&gt;

&lt;p&gt;Threshold routing sends every extraction above a calibrated confidence threshold straight through, and everything below it to a person. This is the operational payoff of calibration: because the score is trustworthy, the threshold becomes a business dial. Tighten it for payments, loosen it for archiving. In anyformat the threshold is set per workflow and the review queue shows the reviewer each uncertain field next to its evidence on the page.&lt;/p&gt;

&lt;p&gt;Try it on your own documents: the free tier gives 50,000 credits with no credit card, enough to run a real sample through a workflow with confidence thresholds on. &lt;a href="https://anyformat.ai/pricing" rel="noopener noreferrer"&gt;Start free&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Validate across fields with deterministic rules
&lt;/h2&gt;

&lt;p&gt;LLMs should not be the last line of defence. Deterministic validation catches hallucinations that look confident: line items must sum to the invoice total, tax identifiers must pass checksum validation (a Spanish NIF has a verifiable control character), dates must fall in plausible ranges, currencies must match the supplier's country. A hallucinated value that passes the model's own confidence check rarely survives arithmetic.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Measure your hallucination rate on your own documents, continuously
&lt;/h2&gt;

&lt;p&gt;A vendor's benchmark tells you how the system performs on the vendor's documents. Your hallucination rate on your documents is the only number that matters. Hold out a labelled evaluation set, re-run it on every model or prompt change, and track three metrics separately: field accuracy, hallucination rate (a wrong value returned confidently) and abstention rate (the system correctly says "not found"). A system that never abstains is hallucinating somewhere. anyformat's Monitoring runs these evaluations per workflow version, so a model swap shows up as a number before it shows up as a payment.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "reducing" hallucinations means in production
&lt;/h2&gt;

&lt;p&gt;Most guides stop at prompting tips and chunking strategies. Those help at the margin. The structural answer is that hallucination risk never reaches zero, so production systems are designed around detection and routing. The pipeline that holds up is the one where a hallucinated value has to get past a calibrated confidence score, a grounding check, a schema constraint and a validation rule, and where the rare one that does costs a human review rather than a wrong payment. None of that should cost you your data: the whole chain runs with zero document retention and, where required, fully air-gapped.&lt;/p&gt;

&lt;p&gt;Related reading on anyformat.ai: &lt;a href="https://anyformat.ai/blog/anyformat-confidence" rel="noopener noreferrer"&gt;what a confidence score means and how we calibrate it&lt;/a&gt; · &lt;a href="https://anyformat.ai/vs/anyformat-vs-llamaparse" rel="noopener noreferrer"&gt;anyformat vs LlamaParse&lt;/a&gt; · &lt;a href="https://anyformat.ai/blog/best-aws-textract-alternatives" rel="noopener noreferrer"&gt;alternatives to AWS Textract&lt;/a&gt; · &lt;a href="https://anyformat.ai/blog/best-google-document-ai-alternatives" rel="noopener noreferrer"&gt;alternatives to Google Document AI&lt;/a&gt; · &lt;a href="https://anyformat.ai/use-cases/api-integration" rel="noopener noreferrer"&gt;API integration&lt;/a&gt; · security and compliance at &lt;a href="https://trust.anyformat.ai" rel="noopener noreferrer"&gt;trust.anyformat.ai&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can LLM hallucinations in document extraction be eliminated completely?&lt;/strong&gt;&lt;br&gt;
No. They can be made rare through grounding, schema constraints and layout-aware parsing, and they can be made detectable through calibrated confidence and deterministic validation, so that undetected errors approach zero even though model errors do not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between confidence and calibrated confidence?&lt;/strong&gt;&lt;br&gt;
Confidence is a number the model emits. Calibrated confidence is a number that has been verified against ground truth so that 95 means 95% correct. Only the second one can safely drive automation decisions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do reasoning models or newer LLMs solve extraction hallucinations?&lt;/strong&gt;&lt;br&gt;
Each model generation shifts the error profile; none eliminates it. The practical consequence is technique 7: re-run your evaluation set on every model change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What hallucination rate is acceptable for invoice processing?&lt;/strong&gt;&lt;br&gt;
The operational target is the undetected error rate after confidence routing and validation, not the model's raw error rate. Teams processing payments set the threshold so that undetected errors are measured in fractions of a percent, and accept the corresponding share of documents that go to review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I know whether a vendor's confidence score is actually calibrated?&lt;/strong&gt;&lt;br&gt;
Ask for the evaluation: the document set, its size, and the accuracy observed inside each confidence band. A calibration claim without a published evaluation is a label.&lt;/p&gt;




&lt;p&gt;Ready to measure it on your documents? &lt;a href="https://anyformat.ai/pricing" rel="noopener noreferrer"&gt;Run a bake-off with anyformat&lt;/a&gt;: your files, your schema, the confidence and evidence on every field.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://anyformat.ai/blog/reduce-llm-hallucinations-document-extraction" rel="noopener noreferrer"&gt;anyformat.ai&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
