DEV Community

Cover image for How to scrape Instagram stories without login (Python, JSON in 30 seconds)
Intropix-team
Intropix-team

Posted on Edited on

How to scrape Instagram stories without login (Python, JSON in 30 seconds)

Instagram stories disappear after 24 hours, which makes them the most annoying content type to work with programmatically. Posts sit there forever and give you time to figure things out. Stories punish you for slow pipelines, meaning if your scraper has a bad day, the data is simply gone.

For a client project I needed active stories from a list of brand accounts, as structured data, on a schedule. This post is the setup I ended up with. It uses no Instagram login, no cookies, no browser automation, and the whole thing is about 15 lines of Python.

Why not just log in with a burner account?

Because Instagram is very good at killing burner accounts. The usual open-source route (instaloader, instagrapi and friends) works right up until it doesn't. First a checkpoint challenge, then forced 2FA, then the session quietly gets shadow-banned, and suddenly you are babysitting a stable of warm accounts and residential proxies instead of doing your actual job. Stories make this worse than posts do: the web UI has no anonymous path to them at all, so every story fetch normally rides on a logged-in session.

The alternative is to let someone else own that problem. I use a hosted scraper on Apify (the Instagram Stories Scraper actor). You send usernames, you get JSON back, and session management stops being your concern. Full disclosure: this is my own actor, so treat me as biased. The technique below is the same with any hosted scraper that exposes an API; the specific field names you'll see are from mine.

What you need

  • A free Apify account (the free tier is enough to test: the actor caps free runs at 10 stories per run, around 20 per day)
  • Your Apify API token, from Console > Settings > API & Integrations
  • Python 3.9+

Step 1: install the client

pip install apify-client
Enter fullscreen mode Exit fullscreen mode

Step 2: run the scraper

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("intropix/instagram-stories-scraper").call(
    run_input={
        "usernames": ["natgeo"],
        "maxResults": 2,
    }
)

items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
print(f"{run['status']}: {len(items)} stories")
for story in items:
    print(story["username"], story["media_type"], story["media_url"][:60])
Enter fullscreen mode Exit fullscreen mode

call() starts the actor and blocks until it finishes. On my last test run this took 27 seconds wall clock including container startup; the scrape itself is 2-3 seconds per account.

What the output looks like

One dataset item per story, always the same 22 fields. This is a real item from the run above (media URL truncated):

{
  "username": "natgeo",
  "user_pk": "787132",
  "full_name": "National Geographic",
  "is_verified": true,
  "is_private": false,
  "story_pk": "3942549043641043661",
  "taken_at": "2026-07-16T13:28:11Z",
  "expiring_at": "2026-07-17T13:28:11Z",
  "media_type": "image",
  "media_url": "https://instagram.frix7-1.fna.fbcdn.net/v/t51.82787-15/747613074_...",
  "video_duration": null,
  "is_paid_partnership": false,
  "attribution_url": null,
  "width": 750,
  "height": 1334,
  "caption": null,
  "link_urls": [],
  "hashtags": [],
  "mentions": ["disneyparks", "disneyland"],
  "music_title": null,
  "music_artist": null,
  "coauthors": []
}
Enter fullscreen mode Exit fullscreen mode

A few fields worth calling out:

  • expiring_at tells you exactly when the story dies, so you know your download deadline.
  • media_url is a direct CDN link to the full-resolution image or video. These links expire after a while, so if you want the file, download it in the same job.
  • is_paid_partnership, mentions and link_urls are the interesting ones for influencer-verification and competitor-monitoring use cases: they tell you who a story tags and where it links.
  • story_pk is a stable numeric ID, use it for dedup.

The schema is fixed. If Instagram adds junk fields upstream, they get dropped rather than leaking into your dataset, so downstream code does not break when Instagram changes something.

Can it read age-restricted (18+) accounts?

Yes, and this was the reason I built it. Alcohol, vape and nightlife brands often mark their accounts age-restricted, and those accounts return nothing to anonymous story viewers and to most scrapers. If you do competitor monitoring in those verticals, or you verify influencer placements for a drinks brand, that gap is exactly where your data needs to be. This scraper reads them like any other public account: I verified delivery against age-gated brand accounts (Budweiser was my test case, 17 story items from an account that an anonymous viewer cannot see at all).

To be clear about scope, this is for public and age-gated brand content. Private accounts stay private and the actor does not touch them.

What does it cost?

Pay-per-event, no subscription. Three events: $0.005 per run start, $0.002 per profile scanned, $0.0025 per story delivered. That lands at "from $2.50 per 1,000 stories", or about $2.70 per 1,000 all-in on typical runs. A daily check on two competitor accounts costs about $0.014 per run. Failed lookups (typo'd or deleted usernames) are never charged.

If you run at serious volume across 100+ accounts, do the math against per-username-priced scrapers, which get cheaper at scale; this one wins on small and mid-size runs.

Running it on a schedule

You could cron the Python script, but Apify has a scheduler built in, so I let the platform do it: actor page > Schedule, pick an interval, done. Add a webhook (Integrations tab) if you want each finished run POSTed to your endpoint.

For no-code pipelines there is an official Apify node for n8n; I published ready-made n8n templates for story archiving to Google Drive, competitor tracking to Sheets with an AI daily brief, and influencer placement verification to Notion. Search "Instagram stories" in the n8n template gallery, or check the examples repo below.

No-dependency Node version

If you would rather not install a client library, the plain REST API does the same job. This endpoint starts a run, waits, and returns the dataset items in one call:

const res = await fetch(
  "https://api.apify.com/v2/acts/intropix~instagram-stories-scraper" +
    "/run-sync-get-dataset-items",
  {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      // Send the token as a header, not in the query string. URLs land in
      // shell history, proxy logs and server access logs.
      Authorization: `Bearer ${process.env.APIFY_TOKEN}`,
    },
    body: JSON.stringify({ usernames: ["natgeo"], maxResults: 2 }),
  }
);
const stories = await res.json();
console.log(stories.length, "stories");
Enter fullscreen mode Exit fullscreen mode

Node 18+ has fetch built in, so that is the entire program.

Getting it into a spreadsheet

The lazy path that works everywhere: flatten to CSV, open in Google Sheets or Excel.

import csv

FIELDS = ["username", "story_pk", "taken_at", "expiring_at", "media_type",
          "caption", "link_urls", "mentions", "is_paid_partnership", "media_url"]

with open("stories.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=FIELDS, extrasaction="ignore")
    writer.writeheader()
    for story in items:
        row = {**story,
               "link_urls": " ".join(story["link_urls"]),
               "mentions": " ".join(story["mentions"])}
        writer.writerow(row)
Enter fullscreen mode Exit fullscreen mode

For a fully automated Sheets pipeline (append rows on every scheduled run), use the n8n template instead of reinventing it.

Wrap-up

Complete runnable versions of all three examples (Python, Node, CSV export) are in the examples repo: github.com/Intropix-team/instagram-scraper-examples. The actor itself, including full field reference and pricing, is at apify.com/intropix/instagram-stories-scraper. The free tier is enough to test whether it fits your use case before any money is involved.

If you need the permanent side of Instagram instead (posts and reels, with like, comment and play counts, follower counts, date-range filtering, and reels that never appear on the profile grid), that is a separate actor, the Instagram Posts & Reels Scraper, and I wrote up that one the same way: How to scrape Instagram posts and engagement data without login.

If you have a stories use case that is close but not quite covered here, leave a comment. I read them, and edge cases are how the actor got most of its fields.

Top comments (0)