DEV Community

Kayvan Zahiri
Kayvan Zahiri

Posted on

Pagination nulls eat your CSV — a 5-check public extract checklist

One-off public extracts usually die in pagination, not in the parser.

Before you ship a CSV from a public HTML/JSON source, run this checklist:

  1. Count pages vs rows — if the site says ~4,200 items and you got 800, you stopped early (or hit a soft wall).
  2. Null spike after page N — plot null rates by page index. A sudden jump usually means the template changed or the bot got a thinner HTML shell.
  3. Stable dedup key — prefer a public id/slug over fuzzy title+city. Re-runs without a key create silent duplicates.
  4. Field cap — keep ≤12 columns. Extra “nice to have” fields are where nested JSON becomes garbage strings.
  5. schema.md — record source pattern, row count, null rates, and what you dropped. Future-you (or a buyer) will thank you.

Prefer official JSON/API when it exists. Flatten nested fields once and note the drop.

Soft offer

I sell a fixed-scope OpsPacket Public Data Pull for public sources only:

  • ≤5,000 rows · ≤12 fields · robots.txt respected
  • CSV + schema.md
  • $149 ≤72h · $199 rush ≤24h

Landing + sample CSV: https://kayvan-zahiri.github.io/opspacket-public-data-pull/

Checkout: https://kayvanandre.gumroad.com/l/public-data-pull · rush https://kayvanandre.gumroad.com/l/public-data-pull-rush

Contact: kayvanandre@gmail.com

Hard no: login walls, CAPTCHA farms, LinkedIn, personal email/phone harvest, ToS-hostile targets.

Top comments (0)