<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sushrut mishra</title>
    <description>The latest articles on DEV Community by sushrut mishra (@sushrut_devable).</description>
    <link>https://dev.to/sushrut_devable</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4086690%2Ff07abe4d-2bbe-42ea-a33e-66e363003921.png</url>
      <title>DEV Community: sushrut mishra</title>
      <link>https://dev.to/sushrut_devable</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sushrut_devable"/>
    <language>en</language>
    <item>
      <title>The best OCR and document extraction APIs for developers in 2026</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Tue, 06 Oct 2026 06:05:25 +0000</pubDate>
      <link>https://dev.to/landingai/the-best-ocr-and-document-extraction-apis-for-developers-in-2026-5156</link>
      <guid>https://dev.to/landingai/the-best-ocr-and-document-extraction-apis-for-developers-in-2026-5156</guid>
      <description>&lt;p&gt;Run a phone photo of a crumpled receipt through a plain OCR engine and you get a wall of characters with the totals in the wrong spots. Useless. That gap, between reading the text and actually understanding the document, is what splits the seven tools here.&lt;/p&gt;

&lt;p&gt;Some you host yourself for free. Some are cloud APIs built for whatever stack you're already on. One reads the layout first and hands back real structure. That's how the list is sorted, and every tool gets a snippet you can paste and run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agentic extraction: reading beyond OCR
&lt;/h2&gt;

&lt;h3&gt;
  
  
  LandingAI's Agentic Document Extraction (ADE)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://landing.ai/agentic-document-extraction" rel="noopener noreferrer"&gt;LandingAI's Agentic Document Extraction (ADE)&lt;/a&gt; isn't really an OCR engine, and that's the point. Plain OCR turns pixels into characters. ADE reads the layout first, works out what each block is, and hands you structured data an agent can just use.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.landing.ai/dpt3/parse" rel="noopener noreferrer"&gt;Parse&lt;/a&gt; reads a document down to its lines and table cells, and grounds every one of them: each value carries its own page, character range, and bounding box, so it points straight back to where it sits on the page. What comes back is reading-order Markdown plus a structured tree of pages, elements, lines, and cells, each with a stable id. Feed that Markdown to &lt;a href="https://docs.landing.ai/ade/ade-extract" rel="noopener noreferrer"&gt;Extract&lt;/a&gt;, and typed JSON comes back with every value still traceable to its page.&lt;/p&gt;

&lt;p&gt;On the infrastructure side, ADE ships Python and TypeScript SDKs and a CLI, plus an &lt;a href="https://docs.landing.ai/dpt3/parse-async" rel="noopener noreferrer"&gt;async Jobs API&lt;/a&gt; for large batches, and a single document can run up to 6,000 pages or 1 GB. Credits run at a &lt;a href="https://docs.landing.ai/ade/ade-pricing" rel="noopener noreferrer"&gt;flat $1 for every 100&lt;/a&gt;, across every tier. For a regulated shop, there's HIPAA-compliant processing with a BAA and a &lt;a href="https://docs.landing.ai/ade/zdr" rel="noopener noreferrer"&gt;Zero Data Retention&lt;/a&gt; option starting at the Team tier, and VPC or on-premises deployment on Enterprise.&lt;/p&gt;

&lt;p&gt;Reach for it when you need structured, checkable output from messy documents, not raw text you'll spend a day reshaping.&lt;/p&gt;

&lt;p&gt;Install with &lt;code&gt;pip install landingai-ade&lt;/code&gt;. A minimal Parse flow, submitting a job and reading back the result, looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;landingai_ade&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LandingAIADE&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LandingAIADE&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads VISION_AGENT_API_KEY
&lt;/span&gt;
&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;document_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://your-storage.example.com/scanned-invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;service_tier&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;standard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;done&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;raise_on_failure&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# done.result.markdown holds the reading-order Markdown, grounded per element
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Managed cloud OCR APIs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Mistral OCR
&lt;/h3&gt;

&lt;p&gt;Mistral OCR is a vision language model that reads a page and gives you Markdown back, with image bounding boxes and structure metadata. It handles 40 plus languages and takes PDFs and the usual image formats, and because the output is Markdown per page, it drops into an LLM pipeline with barely any cleanup. It's a hosted call to Mistral, and it's one of the cheaper managed options per page.&lt;/p&gt;

&lt;p&gt;Reach for it when you've got multilingual docs and you care more about clean Markdown than where it runs. Install with &lt;code&gt;pip install mistralai&lt;/code&gt;. One call, page Markdown out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;mistralai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Mistral&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Mistral&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;ocr_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ocr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;process&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mistral-ocr-latest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;document_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;document_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://example.com/doc.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ocr_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pages&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  AWS Textract
&lt;/h3&gt;

&lt;p&gt;If you're already on AWS, Textract is the path of least resistance. Its &lt;code&gt;detect_document_text&lt;/code&gt; call reads printed and handwritten text and gives you lines and words with coordinates, and when you need forms and tables too, &lt;code&gt;analyze_document&lt;/code&gt; covers that at a higher price. Pricing runs per feature and per page, so a plain text call costs less than a forms-and-tables one.&lt;/p&gt;

&lt;p&gt;Reach for it when you want dependable text plus location data inside an AWS pipeline. It runs through boto3, and the pure OCR call is &lt;code&gt;detect_document_text&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;textract&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scanned-form.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;detect_document_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bytes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()})&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Blocks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BlockType&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LINE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Azure AI Document Intelligence
&lt;/h3&gt;

&lt;p&gt;Azure AI Document Intelligence is the old Form Recognizer, renamed. Its &lt;code&gt;prebuilt-read&lt;/code&gt; model pulls printed and handwritten text out of PDFs and scans, and it reads Word, Excel, PowerPoint, and HTML too, tracking paragraphs, lines, words, and languages along the way. It runs as a cloud call or as an on-premises Docker container, which matters when your data can't leave the building, and the same Read engine underlies Azure's Layout, Invoice, and Receipt models, so moving up to structure later is a small step rather than a rebuild.&lt;/p&gt;

&lt;p&gt;Reach for it when you're on Azure and want OCR with a clear path to prebuilt document models. Install with &lt;code&gt;pip install azure-ai-documentintelligence&lt;/code&gt;. Point &lt;code&gt;prebuilt-read&lt;/code&gt; at a file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;azure.ai.documentintelligence&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DocumentIntelligenceClient&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;azure.core.credentials&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AzureKeyCredential&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;DocumentIntelligenceClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://&amp;lt;resource&amp;gt;.cognitiveservices.azure.com/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;credential&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;AzureKeyCredential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scanned-doc.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;poller&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;begin_analyze_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prebuilt-read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;poller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;result&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Google Cloud Vision
&lt;/h3&gt;

&lt;p&gt;Google Cloud Vision is Google's general image API, and OCR is one feature among many. &lt;code&gt;DOCUMENT_TEXT_DETECTION&lt;/code&gt; handles dense and handwritten pages and returns them as pages, blocks, paragraphs, words, and symbols with bounding boxes, and it figures out the language on its own. Pricing runs per image and per feature, with the first 1,000 units a month free and a flat rate after that.&lt;/p&gt;

&lt;p&gt;Reach for it when you're on GCP and want reliable OCR with plenty of per-word metadata to build on. Install with &lt;code&gt;pip install google-cloud-vision&lt;/code&gt;. Document mode handles dense and handwritten pages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google.cloud&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;vision&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ImageAnnotatorClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;handwritten-note.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;image&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;document_text_detection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_text_annotation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Open source OCR engines
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Tesseract OCR
&lt;/h3&gt;

&lt;p&gt;Tesseract is the old reliable: open source, Apache 2.0, backed by Google, and able to read 100 plus languages entirely on your own hardware. You get plain text out of the box, with hOCR and TSV available if you want word positions. It nails clean scans. Throw handwriting, tables, or a bad photo at it, though, and you're writing the cleanup yourself.&lt;/p&gt;

&lt;p&gt;Reach for it when you want free, self-hosted OCR and your documents are reasonably tidy. Install the Tesseract binary, then &lt;code&gt;pip install pytesseract pillow&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pytesseract&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;PIL&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Image&lt;/span&gt;

&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pytesseract&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;image_to_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scan.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  PaddleOCR
&lt;/h3&gt;

&lt;p&gt;PaddleOCR is Baidu's open source OCR toolkit, also Apache 2.0. It turns PDFs and images into structured, LLM-ready JSON or Markdown, and it reads a lot of languages. PP-OCR pairs detection with recognition, and the PP-Structure pipeline adds layout, tables, and reading order on top. It runs locally on CPU or GPU, which makes it a fit for self-hosted and air-gapped setups.&lt;/p&gt;

&lt;p&gt;Reach for it when you want open source OCR that keeps structure, especially for multilingual or Chinese text. Install PaddlePaddle, then &lt;code&gt;pip install paddleocr&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;paddleocr&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;PaddleOCR&lt;/span&gt;

&lt;span class="n"&gt;ocr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PaddleOCR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ocr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scan.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How the seven compare
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Text type&lt;/th&gt;
&lt;th&gt;Languages&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th&gt;Cost model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://landing.ai/agentic-document-extraction" rel="noopener noreferrer"&gt;LandingAI ADE&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Print, handwriting, scans&lt;/td&gt;
&lt;td&gt;Broad, incl. non-Latin&lt;/td&gt;
&lt;td&gt;Markdown + structured JSON&lt;/td&gt;
&lt;td&gt;Cloud, VPC, on-prem&lt;/td&gt;
&lt;td&gt;Credit-based&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mistral OCR&lt;/td&gt;
&lt;td&gt;Print, scans&lt;/td&gt;
&lt;td&gt;40+&lt;/td&gt;
&lt;td&gt;Markdown&lt;/td&gt;
&lt;td&gt;Managed cloud&lt;/td&gt;
&lt;td&gt;Per page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Textract&lt;/td&gt;
&lt;td&gt;Print, handwriting&lt;/td&gt;
&lt;td&gt;Broad&lt;/td&gt;
&lt;td&gt;Text + blocks&lt;/td&gt;
&lt;td&gt;AWS cloud&lt;/td&gt;
&lt;td&gt;Per feature, per page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure AI Document Intelligence&lt;/td&gt;
&lt;td&gt;Print, handwriting&lt;/td&gt;
&lt;td&gt;Broad&lt;/td&gt;
&lt;td&gt;Text + layout&lt;/td&gt;
&lt;td&gt;Cloud, on-prem container&lt;/td&gt;
&lt;td&gt;Per page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Cloud Vision&lt;/td&gt;
&lt;td&gt;Print, handwriting&lt;/td&gt;
&lt;td&gt;Broad&lt;/td&gt;
&lt;td&gt;Text + blocks&lt;/td&gt;
&lt;td&gt;Cloud, on-prem option&lt;/td&gt;
&lt;td&gt;Per image, per feature&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tesseract OCR&lt;/td&gt;
&lt;td&gt;Print, clean scans&lt;/td&gt;
&lt;td&gt;100+&lt;/td&gt;
&lt;td&gt;Plain text&lt;/td&gt;
&lt;td&gt;Self-hosted&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PaddleOCR&lt;/td&gt;
&lt;td&gt;Print, scans&lt;/td&gt;
&lt;td&gt;Many&lt;/td&gt;
&lt;td&gt;JSON, Markdown&lt;/td&gt;
&lt;td&gt;Self-hosted&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Matching the tool to the job
&lt;/h2&gt;

&lt;p&gt;The right pick follows from what you're actually optimizing for. If you're pulling structured data out of messy or regulated documents, LandingAI ADE is the one built for it: you get grounded structure instead of raw text to reshape yourself. If the job is clean Markdown from multilingual pages, Mistral OCR does that with the least friction.&lt;/p&gt;

&lt;p&gt;If you're running OCR inside a cloud you're already committed to, stay there: Textract on AWS, Azure AI Document Intelligence on Azure, Cloud Vision on GCP. And if the requirement is free, self-hosted, or air-gapped, Tesseract handles clean scans well, while PaddleOCR is the better choice once you also need layout and tables preserved.&lt;/p&gt;

&lt;p&gt;Whatever you land on, test it against your worst documents, not the clean sample in the docs. The photographed receipt, the multi-column scan, the handwritten form: that's where these tools actually fall apart, and it's the part that bites you later if you skip it now.&lt;/p&gt;

&lt;h2&gt;
  
  
  OCR versus document extraction, and a few practical differences
&lt;/h2&gt;

&lt;p&gt;OCR turns pixels into characters. Extraction goes further and reads layout, tables, and fields, then hands back structured data. ADE sits on the extraction side of that line; Tesseract and Cloud Vision sit closer to plain OCR.&lt;/p&gt;

&lt;p&gt;Handwriting support varies more than people expect. Textract, Azure AI Document Intelligence, and Cloud Vision all read it reasonably well. Tesseract is built for printed text and does a poor job on handwriting.&lt;/p&gt;

&lt;p&gt;Running fully offline is possible with a few of these. Tesseract and PaddleOCR run entirely on your own hardware. Azure and Google both offer on-premises container options, and ADE offers VPC and on-premises deployment on its Enterprise tier.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>landingai</category>
      <category>ocr</category>
    </item>
    <item>
      <title>7 document extraction APIs worth building on in 2026</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Tue, 06 Oct 2026 06:04:58 +0000</pubDate>
      <link>https://dev.to/landingai/7-document-extraction-apis-worth-building-on-in-2026-h6e</link>
      <guid>https://dev.to/landingai/7-document-extraction-apis-worth-building-on-in-2026-h6e</guid>
      <description>&lt;p&gt;Picking a document extraction API in 2026 comes down to one question: can you trust the output enough to act on it without someone re-reading the page? Here's the short version, before we get into the details:&lt;/p&gt;

&lt;p&gt;For extraction you can verify value by value, &lt;a href="https://landing.ai/agentic-document-extraction" rel="noopener noreferrer"&gt;LandingAI's Agentic Document Extraction (ADE)&lt;/a&gt; grounds every field to the exact line it came from, and Reducto is its closest agentic peer. If your data has to stay inside AWS, Google Cloud, or Azure, the native service is the easy path. And if budget rules and you can self host, Docling and Unstructured do the job for free.&lt;/p&gt;

&lt;p&gt;Every tool below comes with a minimal snippet, so you can see how it is actually called rather than only what its pricing page claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 7 document extraction APIs, compared
&lt;/h2&gt;

&lt;p&gt;Each option is judged on the four things that decide whether an extraction survives production: how it structures its output, whether it grounds each value back to the page, where it can run, and how it prices at volume.&lt;/p&gt;

&lt;h3&gt;
  
  
  LandingAI's Agentic Document Extraction (ADE)
&lt;/h3&gt;

&lt;p&gt;LandingAI's Agentic Document Extraction (ADE) is built around one idea: ground everything. Every element Parse returns comes with its own page, character range, and bounding box, so you can trace any extracted value back to exactly where it sits on the page. The second generation runs on the &lt;a href="https://docs.landing.ai/dpt3/parse-models" rel="noopener noreferrer"&gt;DPT-3 model family&lt;/a&gt;, and the whole output is built for code to consume, not for a person to skim.&lt;/p&gt;

&lt;p&gt;Parse itself returns reading-order Markdown plus a hierarchical breakdown: pages, elements, lines, table cells, each carrying its own stable id and its grounding. Hand that Markdown to Extract, and it pulls your fields into typed JSON where every value still traces back to its exact page location.&lt;/p&gt;

&lt;p&gt;On the infrastructure side:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python and TypeScript SDKs, plus a CLI&lt;/li&gt;
&lt;li&gt;An &lt;a href="https://docs.landing.ai/dpt3/parse-async" rel="noopener noreferrer"&gt;async Jobs API&lt;/a&gt; built for large batches&lt;/li&gt;
&lt;li&gt;Single documents up to 6,000 pages or 1 GB&lt;/li&gt;
&lt;li&gt;HIPAA-compliant processing with a BAA and a &lt;a href="https://docs.landing.ai/ade/zdr" rel="noopener noreferrer"&gt;Zero Data Retention&lt;/a&gt; option, starting at the Team tier&lt;/li&gt;
&lt;li&gt;SLAs, uptime guarantees, priority rate limits, and VPC or on-premises deployment on Enterprise&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; developers who need to verify extraction field by field, especially in audit and regulated workflows.&lt;/p&gt;

&lt;p&gt;Install it with &lt;code&gt;pip install landingai-ade&lt;/code&gt;. Here's a minimal Parse flow: submit a job, wait for it, read back the result.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;landingai_ade&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LandingAIADE&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LandingAIADE&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads VISION_AGENT_API_KEY
&lt;/span&gt;
&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;document_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://your-storage.example.com/invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;service_tier&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;standard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;done&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;raise_on_failure&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# done.result.markdown holds the reading-order Markdown, grounded per element
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From there, Extract takes that Markdown plus a schema you define and returns typed JSON, with every value still traceable back to its page location. The &lt;a href="https://docs.landing.ai/ade/ade-extract" rel="noopener noreferrer"&gt;Extract API reference&lt;/a&gt; has the exact schema and request format if you want to go deeper.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reducto
&lt;/h3&gt;

&lt;p&gt;Reducto takes on the full document lifecycle in one platform: parse, extract, classify, split, and edit. It grounds every value to a bounding box citation, and it leans on multi-step layout analysis plus a self-correcting OCR pass to handle the pages other tools choke on.&lt;/p&gt;

&lt;p&gt;It returns layout-aware chunks that are ready for retrieval, so you skip a second re-chunking pass.&lt;/p&gt;

&lt;p&gt;It reports processing over four billion pages. Deployment options run cloud, VPC, on-premises, and air-gapped, and higher tiers add SOC 2 Type II, HIPAA, and Zero Data Retention.&lt;/p&gt;

&lt;p&gt;On the independently published LongExtractBench, it completed all 225 long documents at 99.6% precision and recall.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; regulated teams that lead their evaluation with hard documents. Trade-off: the advanced agentic modes cost more credits than basic OCR.&lt;/p&gt;

&lt;p&gt;Install with &lt;code&gt;pip install reductoai&lt;/code&gt;. You upload a file, then run parse or extract against it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;reducto&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Reducto&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Reducto&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads REDUCTO_API_KEY
&lt;/span&gt;&lt;span class="n"&gt;upload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upload&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;extract&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;upload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;file_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;document_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount_due&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  LlamaParse
&lt;/h3&gt;

&lt;p&gt;LlamaParse is the parsing layer inside the LlamaIndex ecosystem, which makes it the obvious pick if you're already building retrieval and agents there. It reads visually complex pages using vision language models and hands back RAG-ready Markdown you can drop straight into a LlamaIndex pipeline.&lt;/p&gt;

&lt;p&gt;It covers 90+ formats and offers a few parse modes, from a fast pass on digital text to a slower, agentic mode reserved for the hardest pages. Because the integration with LlamaIndex is native, you skip the glue code most RAG projects end up writing by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; RAG projects already standardized on LlamaIndex. Trade-off: credits are tiered, and which mode you pick affects the bill more than page count does.&lt;/p&gt;

&lt;p&gt;Install with &lt;code&gt;pip install llama-cloud-services&lt;/code&gt;. A minimal parse returns Markdown you can feed straight into a pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;llama_cloud_services&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LlamaParse&lt;/span&gt;

&lt;span class="n"&gt;parser&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LlamaParse&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads LLAMA_CLOUD_API_KEY
&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;parser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Docling
&lt;/h3&gt;

&lt;p&gt;Docling comes out of IBM Research and now lives under the Linux Foundation as an MIT-licensed open source project. If you want strong layout-aware parsing but need to run it yourself, this is the strongest free option available. It converts PDFs, Office files, HTML, and images into one unified, traceable representation, entirely on your own infrastructure.&lt;/p&gt;

&lt;p&gt;It ships as a Python library, a CLI, an HTTP service, and an MCP server, which makes it a natural fit for air-gapped environments. There's no per-page API cost either, which matters if you're ingesting at high volume on a tight budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; teams comfortable owning their own infrastructure. Trade-off: you're running and tuning the pipeline yourself, with no vendor SLA to fall back on.&lt;/p&gt;

&lt;p&gt;Install with &lt;code&gt;pip install docling&lt;/code&gt;. It converts a file locally and exports Markdown:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;docling.document_converter&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DocumentConverter&lt;/span&gt;

&lt;span class="n"&gt;converter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;DocumentConverter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;converter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export_to_markdown&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Unstructured
&lt;/h3&gt;

&lt;p&gt;Unstructured pairs an open source library with a hosted platform, and it's built for breadth. If you need to pull in many file types rather than squeeze maximum accuracy out of one, this is the tool. It partitions PDFs, emails, HTML, and Office files into typed, metadata-rich elements ready for chunking.&lt;/p&gt;

&lt;p&gt;Connector support is broad on both ends, sources and destinations, which makes it easy to pull mixed document sets into a vector store. It also exposes element-level coordinates, and the hosted platform adds SOC 2 Type II, HIPAA, and in-VPC deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; large-scale, multi-format RAG ingestion. Trade-off: table fidelity and reading order on complex, multi-column layouts need more tuning here than on the vision-first platforms.&lt;/p&gt;

&lt;p&gt;Install with &lt;code&gt;pip install "unstructured[pdf]"&lt;/code&gt;. One call partitions a file into typed elements:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;unstructured.partition.auto&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;partition&lt;/span&gt;

&lt;span class="n"&gt;elements&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;partition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;el&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;elements&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  AWS Textract
&lt;/h3&gt;

&lt;p&gt;AWS Textract is Amazon's managed OCR and form extraction service, and if your stack already lives on AWS, it's the path of least resistance. It reads printed and handwritten text, pulls key-value pairs and tables, and returns geometry coordinates for every block it detects.&lt;/p&gt;

&lt;p&gt;Pricing runs per feature and per page, across Detect Text, Analyze Tables, Analyze Forms, and Analyze Queries. It's tightly wired into the rest of the AWS stack too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; high-volume, AWS-native pipelines with fairly predictable documents. Trade-off: the output is raw, so restoring reading order and semantic structure is work you write yourself.&lt;/p&gt;

&lt;p&gt;It runs through boto3, and it returns blocks rather than your schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;textract&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;analyze_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bytes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()},&lt;/span&gt;
        &lt;span class="n"&gt;FeatureTypes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FORMS&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TABLES&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# resp["Blocks"] holds LINE, KEY_VALUE_SET, TABLE, and CELL blocks
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Turning those blocks into your target JSON is a step you write yourself, often with a second model call.&lt;/p&gt;

&lt;h3&gt;
  
  
  Google Document AI
&lt;/h3&gt;

&lt;p&gt;Google Document AI is Google Cloud's document processing suite, organized around processors for OCR, form parsing, and a handful of specialized document types. If you're standardized on GCP, this is the natural choice. Each processor returns structured fields with bounding regions, and the specialized parsers handle common formats like invoices and receipts right out of the box.&lt;/p&gt;

&lt;p&gt;Pricing runs per processor and per page, depending on which ones a given workflow calls. It wires directly into BigQuery, Vertex AI, and the rest of the Google stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; managed extraction inside GCP. Trade-off: it stays inside its home cloud, so your data stays resident in GCP too.&lt;/p&gt;

&lt;p&gt;You call a processor you have created in the console:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google.cloud&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;documentai&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;documentai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DocumentProcessorServiceClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;processor_path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PROJECT_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PROCESSOR_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;documentai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;RawDocument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;mime_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;process_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;documentai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ProcessRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;raw_document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How the 7 stack up at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Grounding&lt;/th&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th&gt;Pricing&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://landing.ai/agentic-document-extraction" rel="noopener noreferrer"&gt;LandingAI ADE&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Markdown + JSON elements&lt;/td&gt;
&lt;td&gt;Element-level (page, character range, bounding box)&lt;/td&gt;
&lt;td&gt;Cloud, VPC, on-prem&lt;/td&gt;
&lt;td&gt;Credit-based, priced per model and tier&lt;/td&gt;
&lt;td&gt;Verifiable, audit-ready extraction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reducto&lt;/td&gt;
&lt;td&gt;Layout-aware chunks&lt;/td&gt;
&lt;td&gt;Bounding box citations&lt;/td&gt;
&lt;td&gt;Cloud, VPC, on-prem, air-gapped&lt;/td&gt;
&lt;td&gt;Credit-based, complexity-tiered&lt;/td&gt;
&lt;td&gt;Hard docs, regulated teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LlamaParse&lt;/td&gt;
&lt;td&gt;RAG-ready Markdown&lt;/td&gt;
&lt;td&gt;Layout bounding boxes&lt;/td&gt;
&lt;td&gt;Managed cloud, enterprise options&lt;/td&gt;
&lt;td&gt;Credit-based, tiered modes&lt;/td&gt;
&lt;td&gt;LlamaIndex RAG projects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Docling&lt;/td&gt;
&lt;td&gt;Unified doc, Markdown&lt;/td&gt;
&lt;td&gt;Partial layout metadata&lt;/td&gt;
&lt;td&gt;Self-hosted, local, air-gapped&lt;/td&gt;
&lt;td&gt;Free, open source&lt;/td&gt;
&lt;td&gt;Self-hosted ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unstructured&lt;/td&gt;
&lt;td&gt;Typed elements&lt;/td&gt;
&lt;td&gt;Element-level coordinates&lt;/td&gt;
&lt;td&gt;Open source, hosted, in VPC&lt;/td&gt;
&lt;td&gt;Free library, usage-based platform&lt;/td&gt;
&lt;td&gt;Broad multi-format ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Textract&lt;/td&gt;
&lt;td&gt;Raw text, tables, forms&lt;/td&gt;
&lt;td&gt;Block geometry, page level&lt;/td&gt;
&lt;td&gt;AWS cloud&lt;/td&gt;
&lt;td&gt;Per feature, per page&lt;/td&gt;
&lt;td&gt;AWS-native pipelines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Document AI&lt;/td&gt;
&lt;td&gt;Structured fields&lt;/td&gt;
&lt;td&gt;Bounding regions&lt;/td&gt;
&lt;td&gt;GCP cloud&lt;/td&gt;
&lt;td&gt;Per processor, per page&lt;/td&gt;
&lt;td&gt;GCP-native workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How to choose the right document extraction API
&lt;/h2&gt;

&lt;p&gt;The right choice comes down to whatever constraint is actually squeezing your project. A few starting points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Need to verify every value for an audit or a human review?&lt;/strong&gt; Look at the agentic grounding platforms: &lt;a href="https://docs.landing.ai/ade/ade-overview" rel="noopener noreferrer"&gt;LandingAI ADE&lt;/a&gt; and Reducto.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Locked into a specific cloud?&lt;/strong&gt; Use that cloud's native service: Textract for AWS, Google Document AI for GCP, Azure Document Intelligence for Azure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Working with a tight budget and comfortable self-hosting?&lt;/strong&gt; Docling and the Unstructured library get you there for free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Already building on LlamaIndex?&lt;/strong&gt; LlamaParse saves you the integration work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whichever way you lean, run your own test before committing. Take a sample of your hardest documents, run it through two or three finalists, and look closely at the grounding, the table fidelity, and the actual cost on pages that resemble production, not the demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which document extraction APIs run on-premises?&lt;/strong&gt;&lt;br&gt;
LandingAI ADE offers VPC and on-premises deployment on its Enterprise tier, and Reducto offers the same on its enterprise tiers. Docling runs fully locally, air-gapped included. Azure Document Intelligence has a container option, while Textract and Google Document AI stay inside their home clouds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which options are free?&lt;/strong&gt;&lt;br&gt;
Docling and the Unstructured library are free to self-host. LandingAI ADE's &lt;a href="https://docs.landing.ai/ade/ade-pricing" rel="noopener noreferrer"&gt;Explore tier&lt;/a&gt; starts you off with 1,000 free credits on a pay-as-you-go basis, and most other managed APIs offer some kind of free allowance to get started too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which APIs ground extractions back to the source page?&lt;/strong&gt;&lt;br&gt;
LandingAI ADE grounds every element to its page location, character range, and bounding box. Reducto grounds to a bounding box citation. LlamaParse and Unstructured both return layout coordinates you can map back to the page yourself.&lt;/p&gt;

&lt;p&gt;Ready to see it on your own documents? &lt;a href="https://landing.ai/agentic-document-extraction" rel="noopener noreferrer"&gt;Get started free&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>ocr</category>
      <category>landingai</category>
      <category>api</category>
    </item>
    <item>
      <title>Turning a multi page financial PDF into clean, traceable JSON with LandingAI ADE</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Fri, 18 Sep 2026 11:48:49 +0000</pubDate>
      <link>https://dev.to/landingai/turning-a-multi-page-financial-pdf-into-clean-traceable-json-with-landingai-ade-4nb4</link>
      <guid>https://dev.to/landingai/turning-a-multi-page-financial-pdf-into-clean-traceable-json-with-landingai-ade-4nb4</guid>
      <description>&lt;p&gt;The hard part of a financial PDF is rarely the text. It is a table that runs across pages, a total that carries to the last page, and footnotes that change what a number means. You want that as typed JSON your pipeline can use, and in anything regulated you also want to prove each value came from a specific spot on the page. This walks through the three call pipeline in LandingAI's Agentic Document Extraction, our ADE platform, that gets you both: Parse turns the document into grounded structure, Extract shapes it into your schema, and Ground ties every value back to its page.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Call&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;What it returns&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Parse&lt;/td&gt;
&lt;td&gt;The PDF or image&lt;/td&gt;
&lt;td&gt;Reading order Markdown with tables as HTML, plus a grounded tree of pages and blocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extract&lt;/td&gt;
&lt;td&gt;The parsed Markdown and a schema&lt;/td&gt;
&lt;td&gt;Your fields as typed JSON, with per field text ranges&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ground&lt;/td&gt;
&lt;td&gt;The extract metadata and the parse structure&lt;/td&gt;
&lt;td&gt;Every value mapped to a page and a bounding box&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;p&gt;Install the library and set your key. The client reads &lt;code&gt;VISION_AGENT_API_KEY&lt;/code&gt; from the environment, and you pass &lt;code&gt;environment="eu"&lt;/code&gt; if your key is an EU key, since keys are region specific.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pip&lt;/span&gt; &lt;span class="n"&gt;install&lt;/span&gt; &lt;span class="n"&gt;landingai&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;ade&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;landingai_ade&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LandingAIADE&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LandingAIADE&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads VISION_AGENT_API_KEY
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 1: Parse the PDF
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;client.v2.parse&lt;/code&gt; sends the document to the DPT-3 model and returns a single response with three parts: markdown, structure, and metadata. The Parse v2 API takes PDFs and images, so convert other formats first. Tables come back as HTML inside the Markdown, which keeps rows, columns, and spanning cells intact rather than flattening them into text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;parsed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;annual-report.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dpt-3-pro-latest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;  &lt;span class="c1"&gt;# optional: limit the parse to specific pages
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# reading-order Markdown; tables render as HTML
&lt;/span&gt;&lt;span class="n"&gt;structure&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;structure&lt;/span&gt;  &lt;span class="c1"&gt;# tree: document -&amp;gt; pages -&amp;gt; blocks, each grounded
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;page_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;billing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_credits&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two production details worth knowing up front. If some pages fail, the request still succeeds with HTTP 206 and lists the bad pages in &lt;code&gt;metadata.failed_pages&lt;/code&gt;, so a single unreadable scan will not sink the whole job. And a synchronous parse is meant for documents you can wait on; for long filings you move to the jobs API, covered at the end, which is built to &lt;a href="https://landing.ai/llms/handling-multi-hundred-page-documents-in-enterprise-workloads" rel="noopener noreferrer"&gt;handle multi hundred page documents&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Walk the structure to find every table
&lt;/h3&gt;

&lt;p&gt;The structure field is a tree. The root is the document, its children are pages, and a page's children are the blocks on it. Every node carries the same base shape: a type, a stable id, a span giving the start and end character offsets in the Markdown, and a grounding object holding the page, the range of characters, and a normalized bounding box. A table node also has children, which are its cells.&lt;/p&gt;

&lt;p&gt;That regular shape means you can walk the tree once and pull out exactly what you need. Here is a short recursive walk that finds every table and reports the page it sits on.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;find_tables&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;found&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;found&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;found&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;found&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;table&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;found&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;child&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;children&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[]:&lt;/span&gt;
        &lt;span class="nf"&gt;find_tables&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;child&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;found&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;found&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;table&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;find_tables&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;structure&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grounding&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;table &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; on page &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;children&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; cells&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because each table knows its own page, a statement that breaks across three pages comes back as three table blocks you can read in order rather than one flattened blob. That is the difference that lets carried totals and continued rows stay meaningful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Extract into a typed schema
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;client.v2.extract&lt;/code&gt; reads the Markdown and returns JSON in the shape of a schema you define, as a Pydantic model, a dict, or a JSON string. Because it works across the whole document at once, a grand total on the final page and its line items on earlier pages land in one object, and footnotes come through as their own field instead of being dropped.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;LineItem&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Row label in the statement&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Amount for the row&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Financials&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;line_items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;LineItem&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;grand_total&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Total that carries to the final page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;footnotes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Footnote text tied to the table&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extract&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Financials&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;extraction&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# typed JSON in the shape of Financials
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By default, fields your schema asks for that the model cannot support are skipped so extraction continues; pass &lt;code&gt;strict=True&lt;/code&gt; to reject those with a 422 instead, which you want in a pipeline that should fail loudly. Alongside extraction, you get &lt;code&gt;extraction_metadata&lt;/code&gt;, where every field carries the value and the character ranges it was quoted from. That metadata is what makes the next step possible, and it is the &lt;a href="https://landing.ai/llms/document-extraction-for-rag-preparing-structured-outputs-for-vector-databases" rel="noopener noreferrer"&gt;clean, structured JSON&lt;/a&gt; downstream systems and vector databases consume, the same output the &lt;a href="https://landing.ai/llms/best-document-parsing-apis-2026" rel="noopener noreferrer"&gt;best document parsing APIs&lt;/a&gt; are judged on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Ground each value to its page
&lt;/h2&gt;

&lt;p&gt;Extract knows which characters each value came from; Ground turns those text positions into visual ones. You send it the &lt;code&gt;extraction_metadata&lt;/code&gt; from Extract and the &lt;code&gt;structure&lt;/code&gt; from Parse, and it returns every field mapped to the block it was quoted from, with a page number and a bounding box. It runs synchronously and consumes no credits.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;grounded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ground&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;extraction_metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;extraction_metadata&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;structure&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;structure&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;grounded&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grounding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each field comes back tied to its exact place on the page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"grand_total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"block_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"table-cell-42"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"table_cell"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"grounding"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"page"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"range"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"start"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"end"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8131&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"box"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"xmin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.71&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ymin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.88&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"xmax"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.86&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ymax"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.90&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One rule keeps this honest: the extraction metadata must come from an Extract run on the Markdown of the same Parse response you pass as structure. Re parsing shifts the character ranges and invalidates an older extraction, so pair them by the &lt;code&gt;doc_id&lt;/code&gt; the extract metadata carries. The box is normalized zero to one on the page, so you can highlight the value on a rendering, crop the region, or attach it as a citation in a review interface. If your organization runs Zero Data Retention, Ground is off by design, and you compute the same overlap client side by matching each field's ranges against each block's &lt;code&gt;grounding.range&lt;/code&gt;. Either way, that grand total on page 4 now carries proof of where it came from, which is the &lt;a href="https://landing.ai/llms/audit-trails-in-document-ai-tracing-extracted-data-back-to-source-pages" rel="noopener noreferrer"&gt;audit trail&lt;/a&gt; a finance or compliance reviewer needs before trusting an automated number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scale it in production
&lt;/h2&gt;

&lt;p&gt;A synchronous call is fine while you build, and it raises a timeout error on documents too large to finish inline. For those, the jobs API mirrors the same shape: &lt;code&gt;client.v2.parse_jobs&lt;/code&gt; and &lt;code&gt;client.v2.extract_jobs&lt;/code&gt; each expose create, get, list, and wait, and a single job handles up to 6,000 pages or one gigabyte per PDF.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;full-10k.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dpt-3-pro-latest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;parsed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# or poll with get() and list()
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For high concurrency, &lt;code&gt;AsyncLandingAIADE&lt;/code&gt; exposes the same calls with await. Wrap the work in the library's error types so a pipeline degrades cleanly rather than crashing: a partial parse returns 206 with &lt;code&gt;failed_pages&lt;/code&gt;, a job that times out or fails raises &lt;code&gt;JobWaitTimeoutError&lt;/code&gt; or &lt;code&gt;JobFailedError&lt;/code&gt;, and any non success status raises an &lt;code&gt;APIStatusError&lt;/code&gt; carrying the status code and response.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you end up with
&lt;/h2&gt;

&lt;p&gt;Three calls take a messy multi page filing to typed JSON where every value knows its page and its box. Parse preserves the structure and grounds each block, Extract shapes the content to your schema, and Ground makes each number verifiable against the source. Try the flow on your own document in the Playground at &lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt;, then move the long files to the jobs API when you take it to production.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>documentation</category>
    </item>
    <item>
      <title>Where general multimodal LLMs break on document extraction, and when you need a dedicated platform</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Fri, 18 Sep 2026 11:45:32 +0000</pubDate>
      <link>https://dev.to/landingai/where-general-multimodal-llms-break-on-document-extraction-and-when-you-need-a-dedicated-platform-2ila</link>
      <guid>https://dev.to/landingai/where-general-multimodal-llms-break-on-document-extraction-and-when-you-need-a-dedicated-platform-2ila</guid>
      <description>&lt;p&gt;Every team building on documents eventually asks whether a general multimodal model like GPT-4o, Gemini, or Claude is enough, or whether the work needs a dedicated extraction platform. We build one of those, LandingAI's Agentic Document Extraction, our ADE platform, so here is the honest comparison, judged on the dimensions that decide whether extraction holds up in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;General multimodal LLM&lt;/th&gt;
&lt;th&gt;LandingAI ADE&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Grounding&lt;/td&gt;
&lt;td&gt;Returns the value as text, with no coordinate tying it to the page&lt;/td&gt;
&lt;td&gt;Bounding box per line with DPT-3 Pro, per word with DPT-3 Verity; Extract cites the exact word&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consistency across runs&lt;/td&gt;
&lt;td&gt;Probabilistic, so table shape and values can drift call to call&lt;/td&gt;
&lt;td&gt;Standardized JSON every run; Verity transcribes deterministically&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Confidence signal&lt;/td&gt;
&lt;td&gt;None per value by default&lt;/td&gt;
&lt;td&gt;A confidence score for every word with Verity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost at volume&lt;/td&gt;
&lt;td&gt;Billed per token of the whole document, on every run&lt;/td&gt;
&lt;td&gt;Billed per character returned; a basic page runs under a cent on Verity Standard&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Can general multimodal LLMs handle production document extraction?
&lt;/h2&gt;

&lt;p&gt;They are strong for exploration and reasoning, and they give way once you need the same structure from many documents with a record you can defend. The specific places they break:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No coordinate: you get the field value back as text, with no box tying it to the page, so nothing downstream can auto verify it against the source&lt;/li&gt;
&lt;li&gt;Drift across runs: the same page can return a different table shape or a reworded value from one call to the next, because the output is probabilistic&lt;/li&gt;
&lt;li&gt;Hallucination on hard tables: on merged cells, multi page tables, or totals that carry between pages, a general model can return a clean looking number that has no basis on the page&lt;/li&gt;
&lt;li&gt;Token cost: you pay per token of the entire document on every run, which climbs fast at volume&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a concrete head to head, see &lt;a href="https://landing.ai/llms/landingai-ade-vs-gemini-document-processing" rel="noopener noreferrer"&gt;ADE versus Gemini document processing&lt;/a&gt; and &lt;a href="https://landing.ai/llms/confidence-scores-vs-visual-grounding-what-each-architecture-actually-tells-you" rel="noopener noreferrer"&gt;what each architecture tells you about grounding&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you get from ADE instead: grounding you can read
&lt;/h2&gt;

&lt;p&gt;Every element ADE returns carries its location. A value comes back with the page it sits on, its character range in the Markdown, and a bounding box on the image, inside a top level structure object. The shape looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"structure"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"page"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"range"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"start"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1180&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"end"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1197&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"box"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"xmin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;412&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ymin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;604&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"xmax"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;588&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ymax"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;631&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that back to front: &lt;code&gt;box&lt;/code&gt; is where the value physically sits on the page, &lt;code&gt;range&lt;/code&gt; is where it sits in the returned Markdown, and &lt;code&gt;page&lt;/code&gt; says which page. That trace is the thing a general model was never built to emit. On top of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DPT-3 Pro grounds to the line: every text line in a block carries its own bounding box, across all block types, and it detects tables, figures, signatures, and marginalia in reading order&lt;/li&gt;
&lt;li&gt;DPT-3 Verity, in public preview, grounds to the word: every word carries a bounding box and a confidence score, and it transcribes digital text and tables deterministically&lt;/li&gt;
&lt;li&gt;Extract turns that grounding into citations, so any field you pull traces back to a specific word on a specific page, which is what an audit or a human review needs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What makes a pipeline agentic, versus IDP and OCR?
&lt;/h2&gt;

&lt;p&gt;An agentic pipeline reads a page's layout and structure first and adapts to a format it has never seen, rather than following a fixed template or a single OCR pass.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Template IDP: coordinates and rules tuned for one layout, so a shifted column or a new vendor format quietly breaks it, which is exactly where &lt;a href="https://landing.ai/llms/what-happens-when-a-document-ai-system-encounters-a-document-it-was-not-trained-on" rel="noopener noreferrer"&gt;a document it was never trained on&lt;/a&gt; fails&lt;/li&gt;
&lt;li&gt;Single pass OCR: characters with no blocks and no grounding, leaving every bit of meaning to you&lt;/li&gt;
&lt;li&gt;Agentic, in ADE: reads layout down to words and table cells, adapts without templates, and grounds every value, which is why rule based and single pass approaches &lt;a href="https://landing.ai/llms/why-document-ai-accuracy-degrades-under-load-and-how-agentic-architecture-prevents-it" rel="noopener noreferrer"&gt;lose accuracy under load&lt;/a&gt; on the hard pages&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to choose
&lt;/h2&gt;

&lt;p&gt;Reach for a general model when you are prototyping, when volume is low, or when the task is reasoning over a document rather than pulling the same fields from many. Reach for ADE when you need structured extraction at volume, a value you can trace back to the page, and a cost you can predict; most serious pipelines use both, a platform to produce grounded data and a model to reason over it once it is trustworthy. Run one of your own documents through ADE at &lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt; and read the boxes and confidence scores it returns.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>LandingAI's Agentic Document Extraction pricing explained: credits, service tiers, and characters</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Wed, 16 Sep 2026 06:38:36 +0000</pubDate>
      <link>https://dev.to/landingai/landingais-agentic-document-extraction-pricing-explained-credits-service-tiers-and-characters-3fb4</link>
      <guid>https://dev.to/landingai/landingais-agentic-document-extraction-pricing-explained-credits-service-tiers-and-characters-3fb4</guid>
      <description>&lt;p&gt;Flat per page pricing charges the same for a dense contract page and a near empty cover sheet, which quietly overcharges any real document mix. We took the opposite approach with LandingAI's Agentic Document Extraction, our ADE platform, and built the cost controls into the product so a mixed workload stays economical without you engineering around the bill. &lt;/p&gt;

&lt;p&gt;This post walks through the pricing model we landed on and the three factors that decide what any job costs. The &lt;a href="https://landing.ai/blog/agentic-document-extraction-pricing-core-concepts" rel="noopener noreferrer"&gt;pricing core concepts guide&lt;/a&gt; holds the full formulas; what follows is the practical version for the developer staring at a billing object and working out how &lt;code&gt;total_credits&lt;/code&gt; got there.&lt;/p&gt;

&lt;h2&gt;
  
  
  How ADE pricing works at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What is the billing unit&lt;/td&gt;
&lt;td&gt;Credits, where one credit is one US cent, so one dollar buys 100 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What sets a job's cost&lt;/td&gt;
&lt;td&gt;Three factors: the service tier, the Parse model, and the characters returned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What is the default tier&lt;/td&gt;
&lt;td&gt;Standard, at about half the Priority rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Where do I see the cost&lt;/td&gt;
&lt;td&gt;A billing object with &lt;code&gt;total_credits&lt;/code&gt; in every API response, mirrored in the Playground Usage table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does the interface change the price&lt;/td&gt;
&lt;td&gt;For the same tier the rate holds across the Playground, the REST API, and the Python and TypeScript SDKs; note the Playground always runs Priority&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Everything runs on credits
&lt;/h2&gt;

&lt;p&gt;ADE bills in a single unit, credits, and one credit equals one US cent, so a dollar buys 100 credits. That value holds on every plan and in every region, and every ADE service draws from one shared balance, so you track one number for the whole product. You buy credits two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Self service, by card at &lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt;, which suits developers and teams testing a workload or running in production without a contract&lt;/li&gt;
&lt;li&gt;Enterprise, through an annual agreement and standard procurement, which suits volume terms and invoicing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both paths spend the same credits at the same rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three factors that set a job's cost
&lt;/h2&gt;

&lt;p&gt;Once credits sit in your balance, three factors decide how many a job spends: the service tier you pick, the Parse model you run, and the number of characters processed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Service tier, the speed you ask for
&lt;/h3&gt;

&lt;p&gt;The service tier is the lever you control most directly, since it sets how fast the result returns, and the same work at a different speed costs a different amount. It maps onto &lt;a href="https://landing.ai/llms/document-ai-latency-what-determines-processing-speed-in-production" rel="noopener noreferrer"&gt;what determines processing speed in production&lt;/a&gt;, so you trade turnaround for cost deliberately. Tier selection applies to both Parse and Extract, and there are two of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Reach for it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Priority&lt;/td&gt;
&lt;td&gt;Seconds to minutes&lt;/td&gt;
&lt;td&gt;Sync or async&lt;/td&gt;
&lt;td&gt;A person or agent is waiting on the result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;Minutes to hours&lt;/td&gt;
&lt;td&gt;Async only, and the default&lt;/td&gt;
&lt;td&gt;Background pipelines and scheduled ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Standard costs about half of Priority. Because there is no synchronous Standard path, a synchronous call always bills at Priority, so high volume production work belongs on the Standard async default where the same job costs roughly half.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model, the Parse model you run
&lt;/h3&gt;

&lt;p&gt;The DPT-3 family gives you two Parse models, and the one you pick moves the price directly. Both bill the same way, a per page rate plus a rate for every 1,000 output characters, so only the numbers differ.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DPT-3 Pro, the default and highest quality model, is the one to reach for on harder pages: handwriting, figures, charts, diagrams, or poor scans&lt;/li&gt;
&lt;li&gt;DPT-3 Verity, the lower latency model built for clean, machine readable content like digital text and tables, currently in preview, is the cheaper path for simpler pages&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model and tier&lt;/th&gt;
&lt;th&gt;Per page&lt;/th&gt;
&lt;th&gt;Per 1,000 output characters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DPT-3 Pro, Priority&lt;/td&gt;
&lt;td&gt;1.0 credit&lt;/td&gt;
&lt;td&gt;0.50 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DPT-3 Pro, Standard&lt;/td&gt;
&lt;td&gt;0.50 credits&lt;/td&gt;
&lt;td&gt;0.25 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DPT-3 Verity, Priority&lt;/td&gt;
&lt;td&gt;0.30 credits&lt;/td&gt;
&lt;td&gt;0.20 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DPT-3 Verity, Standard&lt;/td&gt;
&lt;td&gt;0.15 credits&lt;/td&gt;
&lt;td&gt;0.10 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things stand out. Standard is half of Priority on every line, and Verity's page rate runs under a third of Pro's. Because the tier and the model both scale the same page, they compound, so the same page on Verity at Standard costs a fraction of what it costs on Pro at Priority. On a typical business document, DPT-3 Pro on Standard runs a median of about 1.5 credits.&lt;/p&gt;

&lt;p&gt;Teams moving from ADE Gen1 typically see 25% to 80% lower per page cost on a mixed workload, and running Verity on Standard for the easy pages is what reaches the top of that range. You avoid &lt;a href="https://landing.ai/llms/the-real-cost-of-building-a-document-extraction-pipeline-in-house" rel="noopener noreferrer"&gt;building that cost optimization yourself&lt;/a&gt;, since the levers sit in the API call rather than in your pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Characters, the content you get back
&lt;/h3&gt;

&lt;p&gt;You pay for the visible content ADE returns, meaning the Markdown text, tables, and fields on the page. Bounding box coordinates, confidence scores, and other structural metadata stay uncounted, so a sparse page returns few characters and costs little while a dense page returns more and costs more, which is what keeps a mixed workload fair. You also hold a direct control here: the &lt;code&gt;pages&lt;/code&gt; parameter skips pages you do not need, and the &lt;code&gt;options&lt;/code&gt; parameter toggles block types on or off, so trimming output you will not use trims the characters you pay for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping simple documents inexpensive
&lt;/h2&gt;

&lt;p&gt;For simple documents that should not cost much, LandingAI's ADE gives you two levers that stack. Run them on DPT-3 Verity, the model built for clean digital text, and keep them on the Standard tier, which costs about half of Priority. Because the charge also tracks characters, a light page of clean text returns little content and lands at the bottom of the range.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every charge is visible in the response
&lt;/h2&gt;

&lt;p&gt;You never have to guess what a job cost, because every ADE API response carries a billing object reporting the credits consumed and the tier it was charged at, and the Playground mirrors the same figures in its downloadable Usage table. Here is the relevant metadata from a real Parse response, a one page invoice run synchronously on DPT-3 Pro:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dpt-3-pro-20260710"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"page_count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"output_markdown_chars"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1779&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"service_tier"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total_credits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.9&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The billing object is the headline: &lt;code&gt;total_credits&lt;/code&gt; is what the job cost, and &lt;code&gt;service_tier&lt;/code&gt; records the rate it paid. The two fields above it produced that number. A synchronous call bills at the Priority rate for DPT-3 Pro, which is 1 credit for the page plus 0.5 credits per 1,000 output characters. The page returned 1,779 characters, so the content charge is about 0.89 credits, and 1 plus 0.89 rounds to the 1.9 you see.&lt;/p&gt;

&lt;p&gt;Now run the same invoice on the cheaper path, DPT-3 Verity on the Standard tier. Standard is async only, so this goes through Parse Jobs, and the billing comes back like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"service_tier"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"standard"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"total_credits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.4&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same document, 0.4 credits against 1.9, a 79% drop. The tier alone roughly halves the bill and the model alone cuts it to about a third; together they compound to that 79%. That is the whole pricing model in one comparison: the document never changed, only the two choices you made in the call.&lt;/p&gt;

&lt;h2&gt;
  
  
  The levers you control
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Default to Standard for background and production work, and pay about half of Priority&lt;/li&gt;
&lt;li&gt;Send clean digital pages to DPT-3 Verity, and reserve DPT-3 Pro for the hard ones&lt;/li&gt;
&lt;li&gt;Trim the response with the &lt;code&gt;pages&lt;/code&gt; and &lt;code&gt;options&lt;/code&gt; parameters so you pay for the characters you actually use&lt;/li&gt;
&lt;li&gt;Remember the Playground always bills at Priority, so read a lower tier's real cost from an API run on Standard or from the per service formula&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where to check the real numbers
&lt;/h2&gt;

&lt;p&gt;The honest way to know what ADE will cost you is to run a representative document through it at &lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt; and read the &lt;code&gt;total_credits&lt;/code&gt; the billing object hands back. Measured per thousand pages, we see ADE Gen2 land at or below common industry benchmarks, and the per service formulas live at &lt;a href="https://docs.landing.ai" rel="noopener noreferrer"&gt;docs.landing.ai&lt;/a&gt; as the reference to trust. If you are weighing ADE on price against the field, our guide to the &lt;a href="https://landing.ai/llms/best-document-parsing-apis-2026" rel="noopener noreferrer"&gt;best document parsing APIs&lt;/a&gt; sets out the comparison, and the full &lt;a href="https://landing.ai/blog/agentic-document-extraction-pricing-core-concepts" rel="noopener noreferrer"&gt;pricing core concepts guide&lt;/a&gt; covers the concepts end to end.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>documentai</category>
    </item>
    <item>
      <title>LandingAI's Agentic Document Extraction, 2nd Generation: new models, agent ready output, and atomic citations</title>
      <dc:creator>sushrut mishra</dc:creator>
      <pubDate>Wed, 16 Sep 2026 06:33:57 +0000</pubDate>
      <link>https://dev.to/landingai/landingais-agentic-document-extraction-2nd-generation-new-models-agent-ready-output-and-atomic-5edc</link>
      <guid>https://dev.to/landingai/landingais-agentic-document-extraction-2nd-generation-new-models-agent-ready-output-and-atomic-5edc</guid>
      <description>&lt;p&gt;LandingAI's Agentic Document Extraction, which we call ADE, now has a second generation, and this post covers what changed for developers building on it. ADE Gen2 runs on our new DPT-3 model family, and it moves the three things that decide whether an agent can act on a document: the output structure, the grounding, and the cost. Our full launch announcement carries the complete detail, and you can read it here: &lt;a href="https://landing.ai/blog/introducing-agentic-document-extraction-gen2" rel="noopener noreferrer"&gt;Introducing Agentic Document Extraction, 2nd Generation&lt;/a&gt;. Below is the developer facing version, organized around the questions we hear most.&lt;/p&gt;

&lt;h2&gt;
  
  
  The release at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;What ADE Gen2 does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Parse models&lt;/td&gt;
&lt;td&gt;DPT-3 family: DPT-3 Pro, and DPT-3 Verity in public preview&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output unit&lt;/td&gt;
&lt;td&gt;Blocks, which replace chunks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structure&lt;/td&gt;
&lt;td&gt;Pages to blocks to lines to words, or to table cells, each block with a stable ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grounding&lt;/td&gt;
&lt;td&gt;Atomic: line level with DPT-3 Pro, word level with DPT-3 Verity; Extract turns it into citations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;APIs&lt;/td&gt;
&lt;td&gt;New v2 Parse and Extract endpoints, plus async Jobs APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parse pricing&lt;/td&gt;
&lt;td&gt;Based on characters returned, rather than pages sent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playground&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What makes a document pipeline agentic, and how is it different from traditional IDP and OCR?
&lt;/h2&gt;

&lt;p&gt;An agentic document pipeline reads a page the way a careful person does and returns structure an agent can act on without a human re reading it first. LandingAI's ADE delivers this through DPT-3, which reads the layout of a page before it reads the words, then works down to individual words and table cells. Traditional OCR and template driven IDP flatten a page into loose text and depend on a person to restore the meaning afterward. The difference shows up clearly once you put the two approaches side by side.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Traditional OCR and template IDP&lt;/th&gt;
&lt;th&gt;LandingAI ADE Gen2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A fixed template per document type&lt;/td&gt;
&lt;td&gt;Reads varied layouts and messy inputs, from handwriting to scans and photos&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flat text output&lt;/td&gt;
&lt;td&gt;Hierarchical blocks with stable IDs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No grounding or page level only&lt;/td&gt;
&lt;td&gt;Atomic grounding to the line or word, tied to the source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A person reads and corrects the output&lt;/td&gt;
&lt;td&gt;An agent consumes the structured output directly&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Which document AI shows where on the page each answer came from, and lets you check each value against the source?
&lt;/h2&gt;

&lt;p&gt;LandingAI's ADE grounds every element atomically, so you can see exactly where on the page each value came from and trace each extracted value back to the source, which is what &lt;a href="https://landing.ai/llms/visual-grounding-and-auditability-how-landingai-ade-makes-every-extraction-defensible" rel="noopener noreferrer"&gt;makes an extraction defensible&lt;/a&gt; in a regulated workflow. Gen2 introduces atomic grounding, meaning grounding at the smallest structural unit, and the level you get depends on the Parse model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DPT-3 Pro grounds to the line: every text line in a block carries its own span pointers and bounding box, across all block types&lt;/li&gt;
&lt;li&gt;DPT-3 Verity grounds to the word: every word carries its own span pointers, bounding box, and a confidence score&lt;/li&gt;
&lt;li&gt;Table cell grounding now adds a bounding box for every cell, more detailed than Gen1&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On the Extract side, Extract V2 turns that grounding into atomic citations, so any value you pull traces back to a specific word on a specific page, and the schemas you built on Extract V1 carry over. Grounding this fine is the foundation for workflows that could not be built before: localized PII and PHI redaction down to the exact word, line, or cell; document comparison that highlights version differences; and human review interfaces where a reviewer edits Markdown in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clean, structured JSON for API first pipelines
&lt;/h2&gt;

&lt;p&gt;ADE Gen2 returns clean, structured JSON built for agents to consume rather than for people to read, which is also what downstream systems and vector databases want when you &lt;a href="https://landing.ai/llms/document-extraction-for-rag-preparing-structured-outputs-for-vector-databases" rel="noopener noreferrer"&gt;prepare structured output for RAG&lt;/a&gt;. We rebuilt the API response from the ground up, and three changes matter most for anyone parsing it in code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chunks are replaced by blocks, and every block carries a stable ID&lt;/li&gt;
&lt;li&gt;A new top level structure object exposes each block's page, character range, bounding box, and atomic grounding, so an agent traverses the response directly&lt;/li&gt;
&lt;li&gt;The Markdown is standardized, so figures, checkboxes, and tables come back the same way on every call, with tables delivered as HTML to preserve structure that plain Markdown would flatten&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your agent reads a predictable shape on every run, which removes a whole class of parsing edge cases you would otherwise defend against by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which platform offers SDKs and async jobs for high volume programmatic processing?
&lt;/h2&gt;

&lt;p&gt;ADE Gen2 ships two client libraries on the v2 APIs, ade-python and ade-typescript, along with a command line ADE CLI, plus async Jobs APIs for high volume programmatic work. Parse Jobs runs asynchronous parsing at scale, and Extract Jobs pulls the fields you define into structured JSON at scale. Both share one job model and one response envelope, so the flow of submitting work and retrieving results stays the same across the platform:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Parse Jobs handles asynchronous parsing of up to 6,000 pages or one gigabyte per PDF&lt;/li&gt;
&lt;li&gt;Extract Jobs handles asynchronous field extraction into structured JSON, on the Priority and Standard tiers&lt;/li&gt;
&lt;li&gt;A unified job model gives you a shared response envelope and job id format across both, with finished jobs signalled by webhook&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Pricing built for a mixed document workload
&lt;/h2&gt;

&lt;p&gt;Parse now charges on the characters it returns rather than the pages you send, so a light page costs the minimum and a dense page costs more, where DPT-2 billed a flat 3 credits per page. Service tiers add a second lever, Priority at 1.0x for when someone is waiting and Standard at 0.5x as the default for pipelines, and model choice adds a third:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DPT-3 Pro, available now, for complex pages: layout, figures, handwriting, scans, non Latin scripts, and math&lt;/li&gt;
&lt;li&gt;DPT-3 Verity, in public preview, for digital text and tables, at roughly 40% of Pro's credits&lt;/li&gt;
&lt;li&gt;Automated routing between the two models, coming by Fall 2026&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Across a mixed workload we project 25% to 80% lower cost, and DPT-3 Verity on the Standard tier brings basic pages to under one cent each. The full model, with the per credit math, sits in our &lt;a href="https://landing.ai/blog/agentic-document-extraction-pricing-core-concepts" rel="noopener noreferrer"&gt;pricing core concepts guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;ADE Gen2 is live today, and the fastest way to see the difference is to run your own documents through it at &lt;a href="https://ade.landing.ai" rel="noopener noreferrer"&gt;ade.landing.ai&lt;/a&gt;. For the complete list of every change in the release, read our full launch announcement: &lt;a href="https://landing.ai/blog/introducing-agentic-document-extraction-gen2" rel="noopener noreferrer"&gt;Introducing Agentic Document Extraction, 2nd Generation&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
