DEV Community

M Saqlain Aslam
M Saqlain Aslam

Posted on

Building a Social Media Downloader: Handling Rate Limits, Format Detection, and Platform API Changes

I run a small tool called SocialSaver - paste a link from Instagram, TikTok, Facebook, or X, and get the video file back. Simple on the surface. Underneath, it's one of the more annoying things I've built, because none of the hard parts are the parts you'd expect.

Here's what actually eats the engineering time.

The stack, briefly

Backend is a NestJS API inside an Nx monorepo; frontend is Angular. Nothing exotic - the complexity isn't in the framework choices, it's in dealing with platforms that actively don't want this to work.

Problem 1: there's no stable API to call

None of these platforms offer a public "give me this video's direct file URL" endpoint for obvious reasons. So extraction means parsing whatever the platform's web client actually loads - embedded JSON blobs in page source, internal GraphQL calls, or CDN URLs buried in a response that's 90% unrelated tracking data.

The catch: that internal structure changes without notice. A parser that worked last week breaks silently this week, and you don't find out until download requests start failing. There's no changelog to subscribe to - the only signal is your own error rate.

Problem 2: rate limiting is adversarial, not polite

This isn't like calling a documented REST API where you get a 429 and a Retry-After header. Platforms vary wildly:

Some return a normal-looking 200 with an empty or garbage payload — so you have to validate response shape, not just status code, to detect throttling
Some rotate which endpoint gets throttled first, so retry logic that assumes "it's always this one path" eventually breaks
IP-based limits mean request volume from a single server can get flagged even with well-behaved traffic

Practical fix that's worked reasonably well: queue + backoff with jitter, and treat "suspiciously empty response" as a rate-limit signal even when the HTTP status says success.

Problem 3: format detection is a moving target

"Download the video" sounds like one job. In practice it's: detect platform → detect content type (single video, carousel, story, reel) → detect available quality/format variants → pick the best one that's actually reachable (some URLs in the response are decoys or region-locked). Each platform encodes this differently, and multi-image carousel posts in particular are where most downloader tools silently fail.

Problem 4: this breaks in production, not in dev

Because none of this is documented, you can't really write forward-looking tests for "what happens when Instagram changes their embedded JSON structure." You find out from real user reports. The practical answer has been: logging that captures the raw unexpected response shape (not just "failed"), so when something breaks, there's enough signal to patch it within hours instead of guessing blind.

Where it's at now

SocialSaver handles Instagram, TikTok, Facebook, and X today - no login, no account required; paste a link and get a file back. The architecture above is basically why a tool that sounds like a weekend project ends up needing ongoing maintenance instead of a one-time build.

Curious if anyone else working on similar scraping/extraction-adjacent tools has found a cleaner way to detect silent format changes than "wait for logs to tell you." Would genuinely like to hear it.

Top comments (0)