Last week I shipped a programmatic SEO site that turns any product name into its HS code — 320 pages, one command, zero manual work.
Today I did the same thing for documents.
The problem
Every B2B team needs to pull structured data out of documents:
- Invoices → line items, tax, totals, vendor, customer
- Receipts → merchant, items, payment method
- Contracts → parties, term, governing law, termination clauses
- Bills of lading → shipper, consignee, containers, ports
- Offer letters → candidate, salary, start date
- Lab reports → test, result, units, reference range
- …and 94 more types
Existing OCR APIs either:
- Cost $0.10+ per document with a $500/mo minimum
- Require you to train a custom model per document type
- Return garbage on layout variations
What I built
Docs-to-JSON — one endpoint, 100+ document types, clean JSON back in ~2 seconds.
Try it: https://boring-saas-infra.pages.dev/document-to-json
Live demo (no signup):
curl -X POST https://document-to-json.boring-saas-infra.workers.dev/extract \
-H "Content-Type: application/json" \
-d '{
"document_type": "invoice",
"text": "INVOICE #INV-2043\nDate: 2026-03-14\nVendor: Acme Supplies Ltd.\n..."
}'
Returns:
{
"document_type": "invoice",
"fields": {
"invoice_number": "INV-2043",
"date": "2026-03-14",
"vendor": "Acme Supplies Ltd.",
"customer": "Brightline Imports Inc.",
"subtotal": 925.00,
"tax": 92.50,
"total": 1017.50,
"currency": "USD",
"line_items": [
{"description": "Copper wire, 2mm", "qty": 120, "unit": "m", "unit_price": 4.50},
{"description": "Steel fittings, 3/4\"", "qty": 40, "unit": "pcs", "unit_price": 7.25}
]
}
}
The distribution play
Same playbook as the HS Codes stream:
-
Product page with live demo —
/document-to-json -
100 SEO landing pages —
/document-type/invoice-to-json,/document-type/contract-to-json,/document-type/lab-report-to-json, etc. - Programmatic generation — one JSON config + one template + one generator script
- Sitemap ping via IndexNow — 426 URLs to Bing/Yandex in one call
The generator
products-docs.json (excerpt):
[
{"slug":"invoice-to-json","name":"Invoice","category":"Finance"},
{"slug":"contract-to-json","name":"Contract","category":"Legal"},
{"slug":"bill-of-lading-to-json","name":"Bill of Lading","category":"Logistics"},
{"slug":"offer-letter-to-json","name":"Offer Letter","category":"HR"},
{"slug":"lab-report-to-json","name":"Lab Report","category":"Medical"}
]
generator-docs.cjs reads the JSON, loops over 100 entries, injects them into a single template, writes 100 folders, updates sitemap.xml and llms.txt.
That's it. One command:
node generator-docs.cjs
Output:
✅ Згенеровано 100 сторінок у output/document-type/
✅ sitemap.xml оновлено — всього 426 URL
✅ llms.txt оновлено — 100 типів документів
SEO structure of each page
Every one of the 100 pages includes:
- Unique title + meta description
- Schema.org
Product+BreadcrumbList+FAQPage - Input sample (raw document text) and output sample (structured JSON)
- cURL example with
document_type=<slug> - Pricing block
- 5 FAQ items specific to that document type
- Related document type chips
Zero duplicate content. Each page has its own example tailored to the document type — invoices get invoice samples, contracts get contract samples, lab reports get lab samples.
Pricing
- Free — 50 extractions/day, all 100+ document types, REST API
- Pro — $19.99/mo, 5,000 extractions, priority queue, email support
Stripe alternative: Suby (crypto + card, instant setup).
Tech stack
- Worker — Cloudflare Workers (already deployed)
- Landing + 100 SEO pages — static HTML on Cloudflare Pages
- Generator — Node.js, no dependencies
-
Deploy —
wrangler pages deploy output - Sitemap ping — IndexNow (Bing + Yandex)
Whole thing ran end-to-end in under 2 hours.
What's next
- Google Sheets add-on:
=DOC2JSON(A1) - MCP server for AI agents to call the API directly
- 3 more programmatic streams: Finance, Monitoring, Compliance
Try it
- Landing: https://boring-saas-infra.pages.dev/document-to-json
- Example SEO page: https://boring-saas-infra.pages.dev/document-type/invoice-to-json
- API: POST https://document-to-json.boring-saas-infra.workers.dev/extract
If you're processing documents at any scale and don't want to pay $$$/mo for OCR — give it a shot. Free tier is genuinely free.
Happy to answer questions about the programmatic SEO setup, the generator, or the worker in comments.
Top comments (0)