DEV Community

Cover image for How to scrape Instagram posts and engagement data without login (Python)
Intropix-team
Intropix-team

Posted on Edited on

How to scrape Instagram posts and engagement data without login (Python)

If you have ever tried to pull post data out of Instagram in code, you already know the official route is a dead end unless you own the account. The Graph API hands you your own business profile and basically nothing beyond it. So anyone doing competitor analysis, influencer verification, or plain content research ends up scraping instead, and then runs straight into the next wall: anonymous scraping of Instagram stopped working years ago.

A client wanted posts and reels from a list of brand accounts, with engagement numbers attached, running on a daily schedule, and able to backfill a historical date range on demand. What I landed on runs without an Instagram login, cookies, or a headless browser, and the core of it is maybe 15 lines of Python.

Why not instaloader and a burner account?

You can, until the account dies. The open-source libraries (instaloader, instagrapi) are genuinely good, but a burner running against them starts collecting checkpoint challenges and shadow-bans within days, and at that point you are running an account-warming operation, not an analytics one. Feeds are a bit friendlier to scrape than stories, but the moment you want reliable engagement counts and deep pagination, Instagram wants to see a logged-in session again.

So I let someone else carry the session problem. I use a hosted scraper on Apify (the Instagram Posts & Reels Scraper actor): usernames in, JSON out. Full disclosure, it's mine, so factor that in. Any hosted scraper with an API gets you the same shape of result; the exact keys below are what this one returns.

What you need

  • A free Apify account (free tier is enough to test: the actor caps free runs at 10 posts per run, around 20 per day)
  • Your Apify API token, from Console > Settings > API & Integrations
  • Python 3.9+

Step 1: install the client

pip install apify-client
Enter fullscreen mode Exit fullscreen mode

Step 2: run the scraper

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("intropix/instagram-posts-reels-scraper").call(
    run_input={
        "usernames": ["natgeo"],
        "maxPosts": 2,
    }
)

items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
print(f"{run['status']}: {len(items)} posts")
for post in items:
    print(post["post_type"], post["like_count"], post["permalink"])
Enter fullscreen mode Exit fullscreen mode

call() blocks until the run finishes. My last test run took 10 seconds wall clock including container startup.

What the output looks like

One dataset item per post, always the same 24 fields, with every media slide nested inside it. This is a real reel from the run above (URLs truncated, caption shortened):

{
  "username": "natgeo",
  "user_pk": "787132",
  "full_name": "National Geographic",
  "is_verified": true,
  "is_private": false,
  "follower_count": 280000000,
  "following_count": 148,
  "profile_pic_url": "https://instagram.fxxx.fna.fbcdn.net/v/t51.2885-19/...",
  "biography": "Experience the world through the eyes of National Geographic photographers.",
  "post_pk": "3512345678901234567",
  "shortcode": "DbrNlfeg60B",
  "permalink": "https://www.instagram.com/p/DbrNlfeg60B/",
  "post_type": "reel",
  "taken_at": "2026-08-05T22:29:25Z",
  "caption": "Black rhinos once again roam Matusadona National Park, decades after poachers turned the Zimbabwean refuge into what one conservationist called...",
  "like_count": 16545,
  "comment_count": 77,
  "view_count": 1253265,
  "is_paid_partnership": false,
  "is_off_grid": false,
  "coauthors": [],
  "hashtags": [],
  "mentions": ["marcuswestbergphotography", "warren_smart"],
  "media": [
    {
      "slide_index": 0,
      "media_type": "video",
      "media_url": "https://scontent-lax3-1.cdninstagram.com/o1/v/t2/f2/m86/...",
      "cover_url": "https://scontent-lax3-2.cdninstagram.com/v/t51.82787-15/...",
      "width": 720,
      "height": 1280,
      "video_duration": 13.696000099182129
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The fields that make this useful for analysis:

  • like_count and comment_count on everything; view_count on reels and videos. Instagram publishes no view count for photos or carousels, so those are null rather than a misleading zero. This is the data the official API will not give you for accounts you don't own.
  • follower_count and following_count: the account's public totals at scrape time, stamped on every item, so a dataset row carries its own denominator for engagement-rate maths.
  • coauthors and is_paid_partnership: collab posts and Instagram's own sponsorship label, the two signals you want for influencer verification.
  • media is one entry per slide. A 10-slide carousel is ONE item with 10 nested slides, not 10 items. Your item counts match reality and so does your bill.
  • post_pk is a stable numeric ID, use it for dedup. permalink gives you the clickable evidence URL.

The schema is a strict allowlist: every item has exactly these 24 keys, every slide the same 7. Upstream changes drop out instead of leaking into your pipeline.

The date filters are the actually good part

sinceDate and untilDate bound the fetch server-side:

run = client.actor("intropix/instagram-posts-reels-scraper").call(
    run_input={
        "usernames": ["natgeo", "bbcearth"],
        "maxPosts": 200,
        "sinceDate": "2026-07-22",
    }
)
Enter fullscreen mode Exit fullscreen mode

That pattern turns a scheduled daily run into an incremental pipeline: yesterday's date in sinceDate, and each run collects exactly the new posts instead of re-downloading (and re-paying for) the whole feed. For one-off historical work, set both bounds and backfill a campaign window post by post.

Two input details worth knowing. maxPosts gives you the N newest posts, which is less obvious than it sounds: Instagram's feed is not date-ordered, pinned posts sit at the top carrying their original dates, and even the tail arrives out of order, so taking the first N the feed hands over quietly returns the wrong posts. The actor reads past the limit, sorts, then truncates. And maxPosts is a total across all usernames, filled account by account in input order, not a per-account allowance. For a fixed number per account, run one account per run (tasks and schedules make that painless) or use date bounds instead.

The reels that aren't on the profile grid

Some reels never show up on the profile grid the feed endpoint reads. Two kinds: reels published only to the reels tab, and trial reels, which Instagram shows to non-followers for a window before the account owner decides whether to keep them. Both are invisible to anything that only walks the grid.

includeOffGridReels (off by default) returns them alongside everything else, and every item carries is_off_grid so you can tell which is which. The flag is on every post, not just when you opt in, so your columns never move.

One thing to know if this is why you're here: trial reels are listed only while the trial is live. When the window closes the reel leaves every public listing without ever joining the grid, so nothing can fetch it afterwards. I watched three of them drop out over about 40 minutes. That makes this a case where a scheduled daily run genuinely beats a backfill, because a backfill cannot recover what is no longer listed.

Can it read age-restricted (18+) accounts?

Yes, same as its stories sibling. Alcohol, vape, OF and nightlife brands often sit behind Instagram's age gate and return nothing to anonymous scrapers. This actor reads their feeds at the standard price and I verified delivery against an age-gated global beer brand's account (carousels and all). Scope boundary: public and age-gated brand content. Private accounts are out.

What does it cost?

Pay-per-event, no subscription: $0.005 per run start, $0.002 per profile scanned, and from $1.90 per 1,000 posts delivered. A carousel counts as one post no matter how many slides. View counts cost nothing extra, and off-grid reels bill as ordinary delivered posts, so turning that flag on spends your maxPosts budget rather than adding a surcharge. Failed lookups (typo'd or deleted usernames) are never charged. A daily incremental pull on two competitor accounts costs about two cents a day; a 1,000-post backfill across 10 accounts lands around $1.93 all-in.

Getting it into a spreadsheet

Flatten to CSV and open it in Sheets or Excel; the engagement columns are what you usually want:

import csv

FIELDS = ["username", "post_pk", "post_type", "taken_at", "permalink",
          "like_count", "comment_count", "view_count",
          "follower_count", "is_off_grid",
          "is_paid_partnership", "hashtags", "mentions"]

with open("posts.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=FIELDS, extrasaction="ignore")
    writer.writeheader()
    for post in items:
        row = {**post,
               "hashtags": " ".join(post["hashtags"]),
               "mentions": " ".join(post["mentions"])}
        writer.writerow(row)
Enter fullscreen mode Exit fullscreen mode

Sort by like_count and you have a poor man's top-content report; group by ISO week of taken_at and you have a posting-cadence benchmark.

No-dependency Node version

The plain REST API does the whole thing in one call:

const res = await fetch(
  "https://api.apify.com/v2/acts/intropix~instagram-posts-reels-scraper" +
    "/run-sync-get-dataset-items",
  {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      // Token in a header, never in the query string: URLs land in
      // shell history, proxy logs and server access logs.
      Authorization: `Bearer ${process.env.APIFY_TOKEN}`,
    },
    body: JSON.stringify({ usernames: ["natgeo"], maxPosts: 2 }),
  }
);
const posts = await res.json();
console.log(posts.length, "posts");
Enter fullscreen mode Exit fullscreen mode

Node 18+ has fetch built in, so that is the entire program.

Running it on a schedule

Apify has a scheduler built in: actor page > Schedule, pick an interval. Pair it with sinceDate for the incremental pattern above, and add a webhook (Integrations tab) if you want each finished run POSTed to your endpoint.

Wrap-up

Complete runnable versions of the examples are in the examples repo: github.com/Intropix-team/instagram-scraper-examples. The actor, with the full 24-field reference and pricing, is at apify.com/intropix/instagram-posts-reels-scraper. If you need the ephemeral side of Instagram instead (stories, which disappear after 24 hours), that is a separate actor with its own write-up: How to scrape Instagram stories without login. The free tier is enough to test whether it fits before any money is involved.

Got a posts or engagement use case that doesn't quite fit the above? Drop it in the comments. The date-range filtering in particular came out of exactly these kinds of questions.

Top comments (0)