DEV Community

Oleksandr
Oleksandr

Posted on

I turned 100 document types into JSON with one command

Last week I shipped a programmatic SEO site that turns any product name into its HS code — 320 pages, one command, zero manual work.

Today I did the same thing for documents.

The problem

Every B2B team needs to pull structured data out of documents:

  • Invoices → line items, tax, totals, vendor, customer
  • Receipts → merchant, items, payment method
  • Contracts → parties, term, governing law, termination clauses
  • Bills of lading → shipper, consignee, containers, ports
  • Offer letters → candidate, salary, start date
  • Lab reports → test, result, units, reference range
  • …and 94 more types

Existing OCR APIs either:

  1. Cost $0.10+ per document with a $500/mo minimum
  2. Require you to train a custom model per document type
  3. Return garbage on layout variations

What I built

Docs-to-JSON — one endpoint, 100+ document types, clean JSON back in ~2 seconds.

Try it: https://boring-saas-infra.pages.dev/document-to-json

Live demo (no signup):

curl -X POST https://document-to-json.boring-saas-infra.workers.dev/extract \
  -H "Content-Type: application/json" \
  -d '{
    "document_type": "invoice",
    "text": "INVOICE #INV-2043\nDate: 2026-03-14\nVendor: Acme Supplies Ltd.\n..."
  }'
Enter fullscreen mode Exit fullscreen mode

Returns:

{
  "document_type": "invoice",
  "fields": {
    "invoice_number": "INV-2043",
    "date": "2026-03-14",
    "vendor": "Acme Supplies Ltd.",
    "customer": "Brightline Imports Inc.",
    "subtotal": 925.00,
    "tax": 92.50,
    "total": 1017.50,
    "currency": "USD",
    "line_items": [
      {"description": "Copper wire, 2mm", "qty": 120, "unit": "m", "unit_price": 4.50},
      {"description": "Steel fittings, 3/4\"", "qty": 40, "unit": "pcs", "unit_price": 7.25}
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

The distribution play

Same playbook as the HS Codes stream:

  1. Product page with live demo — /document-to-json
  2. 100 SEO landing pages — /document-type/invoice-to-json, /document-type/contract-to-json, /document-type/lab-report-to-json, etc.
  3. Programmatic generation — one JSON config + one template + one generator script
  4. Sitemap ping via IndexNow — 426 URLs to Bing/Yandex in one call

The generator

products-docs.json (excerpt):

[
  {"slug":"invoice-to-json","name":"Invoice","category":"Finance"},
  {"slug":"contract-to-json","name":"Contract","category":"Legal"},
  {"slug":"bill-of-lading-to-json","name":"Bill of Lading","category":"Logistics"},
  {"slug":"offer-letter-to-json","name":"Offer Letter","category":"HR"},
  {"slug":"lab-report-to-json","name":"Lab Report","category":"Medical"}
]
Enter fullscreen mode Exit fullscreen mode

generator-docs.cjs reads the JSON, loops over 100 entries, injects them into a single template, writes 100 folders, updates sitemap.xml and llms.txt.

That's it. One command:

node generator-docs.cjs
Enter fullscreen mode Exit fullscreen mode

Output:

✅ Згенеровано 100 сторінок у output/document-type/
✅ sitemap.xml оновлено — всього 426 URL
✅ llms.txt оновлено — 100 типів документів
Enter fullscreen mode Exit fullscreen mode

SEO structure of each page

Every one of the 100 pages includes:

  • Unique title + meta description
  • Schema.org Product + BreadcrumbList + FAQPage
  • Input sample (raw document text) and output sample (structured JSON)
  • cURL example with document_type=<slug>
  • Pricing block
  • 5 FAQ items specific to that document type
  • Related document type chips

Zero duplicate content. Each page has its own example tailored to the document type — invoices get invoice samples, contracts get contract samples, lab reports get lab samples.

Pricing

  • Free — 50 extractions/day, all 100+ document types, REST API
  • Pro — $19.99/mo, 5,000 extractions, priority queue, email support

Stripe alternative: Suby (crypto + card, instant setup).

Tech stack

  • Worker — Cloudflare Workers (already deployed)
  • Landing + 100 SEO pages — static HTML on Cloudflare Pages
  • Generator — Node.js, no dependencies
  • Deploy — wrangler pages deploy output
  • Sitemap ping — IndexNow (Bing + Yandex)

Whole thing ran end-to-end in under 2 hours.

What's next

  • Google Sheets add-on: =DOC2JSON(A1)
  • MCP server for AI agents to call the API directly
  • 3 more programmatic streams: Finance, Monitoring, Compliance

Try it

If you're processing documents at any scale and don't want to pay $$$/mo for OCR — give it a shot. Free tier is genuinely free.

Happy to answer questions about the programmatic SEO setup, the generator, or the worker in comments.

Top comments (0)