Instagram doesn't have a public API for the kind of data OSINT work actually needs —
comment threads with reply hierarchies, full follower lists, carousel slides at original
resolution. So most tools in this space either reverse-engineer the private API (fragile)
or drive a browser (slow, easy to get wrong).
ig-harvester takes the second route,
but with a twist that solves the hardest part of browser-based collection: it attaches to a
Chrome window you're already logged into via the Chrome DevTools Protocol (CDP).
Disclosure: I built ig-harvester. It's MIT-licensed and free. This post explains how it
works, not just what it does.
The core idea: bring your own session
Most scrapers try to become logged in — storing cookies, solving challenges, rotating
accounts. That's a losing arms race, and it puts the tool in a bad spot legally and
operationally.
ig-harvester instead launches Chrome with remote debugging (one PowerShell script), you log
in with a burner account in that window, and the tool attaches to it:
.\dependencies\launch-chrome-debug.ps1 # Chrome with CDP on :9222
# log into your burner account, keep the window open
.\ig-harvester.ps1 scrape someuser --posts 50 --comments --followers --sqlite
Everything the harvest does, your session could do by hand. No API keys, no token storage,
nothing leaves your machine.
Triple-source extraction
The single biggest failure mode of Instagram scrapers is relying on one signal.
Instagram serves content three different ways, and each one breaks at different times:
-
Embedded Relay JSON — the
__ NEXT_DATA__-style blob in the page source. Rich (carousels, comment threads, exact timestamps) but occasionally absent. - Semantic DOM — walking the accessibility-oriented markup. Always present, but shallower.
- Screenshots — the ground truth for anything the first two disagree on.
ig-harvester mines all three and cross-checks. When live Instagram stopped embedding sidecar
JSON for carousels (it did, recently), the tool didn't break — it fell back to walking the
post with the Next button and reading each slide's <img>/srcset at full resolution.
Resumability is not optional
A 500-post harvest takes a while, and browsers crash. Every run writes a SQLite cache keyed
by shortcode:
out/someuser/
someuser.json # payload + analytics
someuser-posts.csv # one row per post
someuser-comments.csv # one row per comment
someuser-followers.csv
someuser-following.csv
someuser.db # --sqlite
.cache.db # resume cache
Kill the process halfway, run it again, and it continues where it stopped instead of
re-collecting the first 250 posts. --no-resume disables it when you want a clean sweep.
Pacing like a human, not a loop
Rate limiting is where "it works on my machine" dies in production. The tool uses:
-
jittered delays between actions (
--delay-min 900 --delay-max 2600, in ms) - a token bucket for burst control
- exponential backoff on transient failures
- optional HTTP/SOCKS5 proxy support for operational security
node dependencies/scrape-ig.mjs --profile someuser --proxy http://127.0.0.1:8080 --delay-min 1200
Data out, insights out
Raw data is the start, not the end. Every run computes engagement rate, posting cadence,
best posting hour/day, top hashtags and mentions, follower/following ratio, and bio signals
(email/phone/URL detection) — straight into the JSON payload.
Then there's ig-analyzer, a live dashboard on localhost:8080 that charts likes per post,
posts by hour/day, and content mix for any scraped or archived profile.
For the archive-minded: import-data.ps1 copies each run into a dated snapshot (re-scrapes
never overwrite history), so follower counts stay comparable over time.
One entry point, no CLI archaeology
.\ig-harvester.ps1 # interactive menu
.\ig-harvester.ps1 scrape someuser --posts 50 --comments
.\ig-harvester.ps1 -Action Full -User someuser -Yes # scrape -> images -> archive
.\ig-harvester.ps1 -Action Doctor # environment check
-Yes makes it non-interactive for CI, -DryRun prints the exact commands, and exit codes
are deterministic (0 ok, 1 step failed, 2 usage error). On macOS/Linux, run the Node
CLIs directly — same flags, same behavior.
Get it
- Repo: github.com/anurag-panda-dev/ig-harvester
- Site & full CLI reference: anurag-panda-dev.github.io/ig-harvester
- Requires Node ≥ 20 and Chrome.
npm install && npm testto verify your setup.
Ethical use: ig-harvester is intended for legitimate OSINT research, journalism,
security audits and education. You are responsible for complying with applicable laws,
Instagram's Terms of Service and privacy regulations (GDPR, CCPA) — and for confirming
consent before collecting data about someone.
Have you built browser-based collectors for sites with no API? I'd love to hear how you
handled extraction fallbacks — the carousel problem was the interesting one for me.
Top comments (0)