DEV Community

Cover image for ig-harvester: an open-source Instagram OSINT toolkit built on Playwright + CDP
Anurag Panda
Anurag Panda

Posted on Originally published at anurag-panda-dev.github.io

ig-harvester: an open-source Instagram OSINT toolkit built on Playwright + CDP

Instagram doesn't have a public API for the kind of data OSINT work actually needs —
comment threads with reply hierarchies, full follower lists, carousel slides at original
resolution. So most tools in this space either reverse-engineer the private API (fragile)
or drive a browser (slow, easy to get wrong).

ig-harvester takes the second route,
but with a twist that solves the hardest part of browser-based collection: it attaches to a
Chrome window you're already logged into
via the Chrome DevTools Protocol (CDP).

Disclosure: I built ig-harvester. It's MIT-licensed and free. This post explains how it
works, not just what it does.

The core idea: bring your own session

Most scrapers try to become logged in — storing cookies, solving challenges, rotating
accounts. That's a losing arms race, and it puts the tool in a bad spot legally and
operationally.

ig-harvester instead launches Chrome with remote debugging (one PowerShell script), you log
in with a burner account in that window, and the tool attaches to it:

.\dependencies\launch-chrome-debug.ps1    # Chrome with CDP on :9222
# log into your burner account, keep the window open
.\ig-harvester.ps1 scrape someuser --posts 50 --comments --followers --sqlite
Enter fullscreen mode Exit fullscreen mode

Everything the harvest does, your session could do by hand. No API keys, no token storage,
nothing leaves your machine.

Triple-source extraction

The single biggest failure mode of Instagram scrapers is relying on one signal.
Instagram serves content three different ways, and each one breaks at different times:

  1. Embedded Relay JSON — the __ NEXT_DATA__-style blob in the page source. Rich (carousels, comment threads, exact timestamps) but occasionally absent.
  2. Semantic DOM — walking the accessibility-oriented markup. Always present, but shallower.
  3. Screenshots — the ground truth for anything the first two disagree on.

ig-harvester mines all three and cross-checks. When live Instagram stopped embedding sidecar
JSON for carousels (it did, recently), the tool didn't break — it fell back to walking the
post
with the Next button and reading each slide's <img>/srcset at full resolution.

Resumability is not optional

A 500-post harvest takes a while, and browsers crash. Every run writes a SQLite cache keyed
by shortcode:

out/someuser/
  someuser.json            # payload + analytics
  someuser-posts.csv       # one row per post
  someuser-comments.csv    # one row per comment
  someuser-followers.csv
  someuser-following.csv
  someuser.db              # --sqlite
  .cache.db                # resume cache
Enter fullscreen mode Exit fullscreen mode

Kill the process halfway, run it again, and it continues where it stopped instead of
re-collecting the first 250 posts. --no-resume disables it when you want a clean sweep.

Pacing like a human, not a loop

Rate limiting is where "it works on my machine" dies in production. The tool uses:

  • jittered delays between actions (--delay-min 900 --delay-max 2600, in ms)
  • a token bucket for burst control
  • exponential backoff on transient failures
  • optional HTTP/SOCKS5 proxy support for operational security
node dependencies/scrape-ig.mjs --profile someuser --proxy http://127.0.0.1:8080 --delay-min 1200
Enter fullscreen mode Exit fullscreen mode

Data out, insights out

Raw data is the start, not the end. Every run computes engagement rate, posting cadence,
best posting hour/day, top hashtags and mentions, follower/following ratio, and bio signals
(email/phone/URL detection) — straight into the JSON payload.

Then there's ig-analyzer, a live dashboard on localhost:8080 that charts likes per post,
posts by hour/day, and content mix for any scraped or archived profile.

For the archive-minded: import-data.ps1 copies each run into a dated snapshot (re-scrapes
never overwrite history), so follower counts stay comparable over time.

One entry point, no CLI archaeology

.\ig-harvester.ps1                                    # interactive menu
.\ig-harvester.ps1 scrape someuser --posts 50 --comments
.\ig-harvester.ps1 -Action Full -User someuser -Yes    # scrape -> images -> archive
.\ig-harvester.ps1 -Action Doctor                      # environment check
Enter fullscreen mode Exit fullscreen mode

-Yes makes it non-interactive for CI, -DryRun prints the exact commands, and exit codes
are deterministic (0 ok, 1 step failed, 2 usage error). On macOS/Linux, run the Node
CLIs directly — same flags, same behavior.

Get it

Ethical use: ig-harvester is intended for legitimate OSINT research, journalism,
security audits and education. You are responsible for complying with applicable laws,
Instagram's Terms of Service and privacy regulations (GDPR, CCPA) — and for confirming
consent before collecting data about someone.


Have you built browser-based collectors for sites with no API? I'd love to hear how you
handled extraction fallbacks — the carousel problem was the interesting one for me.

Top comments (0)