DEV Community

Kayvan Zahiri
Kayvan Zahiri

Posted on Originally published at kayvan-zahiri.github.io

Before you build a scraper: a public-data CSV scope card that keeps one-offs honest

Most “quick scrapes” fail the same way: unclear fields, no row cap, and silent nulls after pagination.

Before you open Playwright (or hire someone), write this scope card for any public source:

  1. URL pattern — exact pages or API endpoints (include pagination / sitemap).
  2. Public-only check — no login wall, no CAPTCHA farm, robots.txt / ToS allow automated fetch.
  3. Fields (≤12) — name each column; drop “nice to have” extras.
  4. Row budget — e.g. ≤5,000 rows. Caps force honesty about coverage.
  5. Dedup key — what makes two rows the same (id, url, …).
  6. Delivery — CSV plus a one-page schema.md with row count, null rates, and the source pattern.

Prefer official JSON/API when it exists (?$limit= on open portals often beats HTML parse). Flatten nested fields once; note what you dropped.

Sample shape

I keep a public Open Brewery DB (California) slice as an example CSV + schema on the landing below so buyers (and future-me) can see the output shape before paying.

Soft offer

I’m the maker of OpsPacket Public Data Pull — fixed-scope public extracts only:

  • ≤5,000 rows · ≤12 fields · robots.txt respected
  • CSV + schema.md
  • $149 within 72 hours · $199 rush within 24 hours

Landing + sample: https://kayvan-zahiri.github.io/opspacket-public-data-pull/

Checkout: https://kayvanandre.gumroad.com/l/public-data-pull · rush https://kayvanandre.gumroad.com/l/public-data-pull-rush

Hard no: login-walled SaaS exports, LinkedIn, personal email/phone harvest lists, ToS-hostile targets. Paste a public URL + fields and I’ll say fit / out-of-scope in one reply.

No ranking, traffic, or revenue guarantees — just a scoped CSV.

Top comments (0)